LF LevelField.io
Sign in Request a demo
Insights Evidence

What a citation actually has to prove before a partner will sign under it

Author Maren EllisCounsel in residence 14 Jul 2026 | 9 min read
Post hero · marked-up brief on a desk

Every vendor in this market quotes a retrieval accuracy number. Almost none of them quote the number a supervising partner actually needs, which is narrower, harder, and considerably less flattering.

The wrong benchmark

Retrieval accuracy asks whether the system found a relevant passage. That is a question about search. Supervision asks something else entirely: whether the proposition in front of me is supported by the passage cited, whether that passage says what the summary claims it says, and whether it was still good law on the day I signed.

Those are three separate failures, and a single accuracy figure hides all of them. A system can retrieve the correct authority and still characterise it wrongly. It can characterise it perfectly and still cite a decision that was reversed eight months ago.

Worked example

In one audit, a competing tool cited the correct Delaware opinion for a change-of-control proposition — but drew the proposition from the losing party's brief, quoted inside the opinion's summary of argument. Retrieval: correct. Answer: wrong.

What supervision needs

A partner reviewing an associate's memo does not re-run the research. They spot-check: pick the two propositions that carry the argument, follow them to source, and judge whether the rest of the work was built the same way. The review is fast because the trail is legible.

That is the standard a machine has to meet. Not "trust the output" — make the output checkable in the same three minutes.

"If I cannot check it faster than I could have written it, it has not saved me anything. It has moved the work."

Managing partner, AmLaw 100 litigation group

Four tests we run

Every release is evaluated against a fixed set of matters with known answers, scored by practising lawyers rather than by string overlap.

01  Support — does the passage carry the claim?98.6% 02  Attribution — is it the holding, not the argument?97.1% 03  Currency — was it good law on the filing date?99.4% 04  Completeness — is contrary authority surfaced?91.2%

What we still get wrong

Completeness is the weakest of the four, and we publish it anyway. Surfacing contrary authority requires a judgment about what a reasonable opponent would raise, and that judgment is where the model is furthest from a good associate.

We would rather show the number and let you supervise accordingly than average it into something comfortable.

Author Maren Ellis Twelve years in commercial litigation before joining LevelField, where she runs the evaluation programme with the engineering team.

Read next

All insights →
Post image Evidence Hallucination is a supervision problem, not a model problem 02 Jul 2026 | 7 min Post image Evidence Reading 1,840 files is not the same as knowing the record 08 May 2026 | 6 min Post image Security What a CISO asks that a legal buyer forgets to 11 Jun 2026 | 8 min