Lexsys

12 March 2026 · 6 min read

Verification, not generation, is the product

Any model can draft a demand letter. The question that decides whether legal AI is usable is how quickly a lawyer can establish that the draft is right.

The first wave of legal AI was sold on generation. Give the system a file and it returns a summary, a chronology, a draft claim. That capability is now close to commoditised: the underlying models are widely available and the gap between the best and the merely adequate has narrowed considerably.

What has not been commoditised is verification. A draft a lawyer cannot check faster than they could have written it is not a saving — it is a transfer of work from drafting to review, and review of unfamiliar text is the slower of the two. The firms getting real leverage out of these systems are the ones that treated verification as the design problem from the start.

What verifiable output looks like

Every factual assertion in a generated document should be traceable to a specific page, clause or field in the source file, and that trace should be one click away rather than a search. When a system states that an agreement carries an effective rate of 24.8 per cent, the lawyer should see the clause it read to reach that figure without leaving the document.

The same applies to what the system did not find. A summary that silently omits a missing signature page is more dangerous than one that flags the gap, because it invites the reader to assume completeness. Absences need to be surfaced as explicitly as findings.

Confidence has to be calibrated, not decorative

Many tools attach confidence scores that turn out to be uniform across easy and hard extractions alike. That is worse than no score: it teaches the reviewer to skim. A score is only useful if low confidence reliably correlates with cases the reviewer would in fact have corrected.

The practical test is simple. Sample a hundred outputs, check them all, and see whether the errors cluster where the system said they might. If they do not, the score is noise and should be removed rather than displayed.

The economics of review

In consumer claims work the unit that matters is minutes per file, end to end, including review. A system that cuts preparation from ninety minutes to five but adds twenty-five minutes of careful checking has delivered a real three-fold gain. A system that cuts preparation to five minutes but leaves the lawyer re-reading the whole agreement to trust the output has delivered nothing at all.

This is why verification design, rather than model choice, is where the difference between tools now sits. The reasoning quality of frontier models is largely shared. The interface that lets a lawyer accept or reject that reasoning in seconds is not.

What to ask a vendor

Ask to see a wrong answer. Ask how the system behaves on a file with a missing page, a scanned page at an angle, and a clause that contradicts another clause. Ask what the reviewer sees, and how long the correction takes. A demonstration on a clean file tells you almost nothing about the tool you will actually be using at volume.

Let's do something great together.

Talk to our team about AI-enabled procurador services, case automation and litigation funding.