How to measure whether legal AI is actually accurate

Date

August 9 2026

mr club

Date

August 9 2026

Accuracy claims are easy to publish and hard to check. A firm evaluating legal AI needs a method that tests what lawyers actually care about: real authorities, correctly cited, supporting the right proposition.

Every legal AI vendor publishes an accuracy figure. Few publish the questions, the scoring rubric, or the things the test deliberately leaves out. That matters, because a legal researcher’s job is not to be impressed by a number; it is to decide whether the answer is right.

This is the method we use to evaluate Themis, and the method any firm can use to evaluate a legal AI product before letting it near a client file.

1. Build a fixed question set

Start with questions that look like real work. In a Mauritian context, that means questions across company law, employment, land, family, civil procedure and financial services regulation. Each question should require a legal proposition supported by one or more authorities.

Avoid two common traps:

  • Trivia questions. A system that can name the sections of an Act is not the same as a system that can apply them.
  • Questions the corpus cannot answer. If the answer depends on unreported facts or a judgment not yet published, the test becomes a test of guessing.

2. Define "correct" tightly

A response is correct only if:

  • every authority it cites exists;
  • each authority supports the proposition it is cited for;
  • the proposition actually answers the question asked.

That definition excludes half-right answers, plausible-sounding hallucinations and answers that cite real cases for the wrong reason.

3. Score citations, not fluency

A well-written wrong answer is worse than a clumsy right one. The scoring rubric should weight the citation check above style. Ask:

  • Did the system provide a verifiable citation?
  • Does the citation say what the system claims it says?
  • Is the citation the best authority for the point, or just a related one?

4. Compare against the alternative

If the claim is that the product outperforms generic AI, test it on the same question set with the same rubric. Use the provider’s documented best practices for prompting. Record the model name, version and date.

5. Test the edges deliberately

Every system is strong in the middle and weak at the edges. A serious evaluation includes:

  • questions that combine several unsettled areas;
  • questions that depend on a recent judgment not yet in the corpus;
  • questions that require factual assumptions the system does not have.

The score on these questions is at least as important as the headline number. A vendor that reports only the easy questions is not telling the whole story.

6. Re-run and publish

A one-off test is marketing. A repeatable test is a methodology. Re-run the evaluation after every material update to the corpus or model, log the results, and publish enough of the method that an independent reviewer can check it.

What we publish

On Themis methodology we set out what we tested, how we scored, what the comparison measures, and what the published accuracy figure deliberately excludes. If a vendor cannot answer those same questions, the accuracy number is not evidence. It is decoration.