12 March 2026 · 6 min read
Verification, not generation, is the product
Any model can draft a demand letter. The question that decides whether legal AI is usable is how quickly a lawyer can establish that the draft is right.
The first wave of legal AI was sold on generation. Give the system a file and it returns a summary, a chronology, a draft claim. That capability is now close to commoditised: the underlying models are widely available and the gap between the best and the merely adequate has narrowed considerably.
What has not been commoditised is verification. A draft a lawyer cannot check faster than they could have written it is not a saving — it is a transfer of work from drafting to review, and review of unfamiliar text is the slower of the two. The firms getting real leverage out of these systems are the ones that treated verification as the design problem from the start.
What verifiable output looks like
Every factual assertion in a generated document should be traceable to a specific page, clause or field in the source file, and that trace should be one click away rather than a search. When a system states that an agreement carries an effective rate of 24.8 per cent, the lawyer should see the clause it read to reach that figure without leaving the document.
The same applies to what the system did not find. A summary that silently omits a missing signature page is more dangerous than one that flags the gap, because it invites the reader to assume completeness. Absences need to be surfaced as explicitly as findings.
Confidence has to be calibrated, not decorative
Many tools attach confidence scores that turn out to be uniform across easy and hard extractions alike. That is worse than no score: it teaches the reviewer to skim. A score is only useful if low confidence reliably correlates with cases the reviewer would in fact have corrected.
The practical test is simple. Sample a hundred outputs, check them all, and see whether the errors cluster where the system said they might. If they do not, the score is noise and should be removed rather than displayed.
The economics of review
In consumer claims work the unit that matters is minutes per file, end to end, including review. A system that cuts preparation from ninety minutes to five but adds twenty-five minutes of careful checking has delivered a real three-fold gain. A system that cuts preparation to five minutes but leaves the lawyer re-reading the whole agreement to trust the output has delivered nothing at all.
This is why verification design, rather than model choice, is where the difference between tools now sits. The reasoning quality of frontier models is largely shared. The interface that lets a lawyer accept or reject that reasoning in seconds is not.
What to ask a vendor
Ask to see a wrong answer. Ask how the system behaves on a file with a missing page, a scanned page at an angle, and a clause that contradicts another clause. Ask what the reviewer sees, and how long the correction takes. A demonstration on a clean file tells you almost nothing about the tool you will actually be using at volume.
Keep reading
27 February 2026
What changes when capacity stops being the constraint
Automating case preparation does not simply let a firm do more of the same. It moves the bottleneck, and the next one is rarely where partners expect.
14 January 2026
How Spanish courts are absorbing automated filings
Procedural infrastructure, not model capability, now sets the pace of legal automation in Spain. What that means for firms filing at volume.
Let's do something great together.
Talk to our team about AI-enabled procurador services, case automation and litigation funding.