Legal Teams// diagnostic

Every clause comes back flagged, so the reviewer stops reading the flags

In short

When a contract review tool flags too many clauses, the model is almost never at fault. A playbook written as ideal sentences flags every acceptable variant of them, and one with no severity tier puts a governing-law nit beside an uncapped indemnity. Grade 50 flags would-have-changed or would-have-accepted; the split names the repair.

Key takeaways

  • Grade 50 flags before touching anything. Wrong flags and unranked flags look identical and need opposite repairs.
  • A playbook written as preferred wording flags every acceptable paraphrase of that wording, which is most of the market.
  • With no severity tier, a governing-law nit and an uncapped indemnity arrive as the same object, so both get skimmed.
  • Semantic equivalence is the largest single bucket: 3 phrasings of one accepted position counted as 3 deviations.
  • Clause types where the firm has no position should say so. Silence dressed as a flag is the worst output available.
  • Precision above roughly 70 per cent turns the problem from calibration into ranking, and the fix changes accordingly.

A reviewer who opens a 40-page agreement, sees 61 flags and closes the panel has made a rational decision. Reading 61 items to find the 4 that matter is slower than reading the contract, and everyone works this out inside a week. The tool is not being ignored because lawyers distrust software. It is being ignored because it has stopped carrying information.

Almost all of this is playbook calibration, and almost none of it is the model. The playbook is a set of written positions; if those positions are expressed as exact sentences rather than acceptable ranges, every reasonable paraphrase in the market becomes a deviation and the flag count tracks document length instead of risk.

50 flags, 1 lawyer, and 2 columns

The measurement comes first because the two failure modes are indistinguishable from the flag count alone. A tool producing 61 useless flags and a tool producing 61 correct flags in no useful order both look like noise to the reviewer, and they need opposite repairs.

  1. Take 50 consecutive flags from real reviews, not from a demo set. Consecutive matters: cherry-picking by clause type reproduces whatever bias created the problem.
  2. Have 1 lawyer mark each flag would-have-changed or would-have-accepted, answering only whether they would have negotiated this wording, with no third option. A maybe column collapses the measurement.
  3. Discard flags raised against text that was not operative before scoring anything. A flag on a superseded clause is a different defect entirely, diagnosed in the review that is right about a superseded version, and counting it as a false positive sends you to the wrong fix.
  4. Record the clause type against every flag. The distribution matters more than the total: 34 of 50 landing on 3 clause types is a playbook-entry problem, while an even spread across 15 types is a threshold problem.
  5. Compute the would-have-changed share. That single number is the precision of the flag stream, and it routes the rest of this page.
  6. Ask the same lawyer for the flags they expected and did not get. Misses never appear in a false-positive count, and a playbook tuned only on precision quietly stops catching things.

What the would-have-changed share tells you to do next

The bands below are working thresholds rather than published standards, and they are useful because each one implies a different piece of work. Set your own after 2 rounds of measurement; what matters is that the number decides the action rather than the other way round.

Would-have-changed shareWhat it usually meansWhere to work first
Below 30 per centThe playbook encodes preferred wording, not acceptable rangesRewrite the 5 highest-volume entries as ranges before anything else
30 to 60 per centReal deviations exist but sit under a pile of paraphrase hitsSemantic equivalence: collapse variants of positions you already accept
60 to 80 per centFlags are broadly right and arrive in no defensible orderAdd severity tiers; stop changing the detection at all
Above 80 per cent, still ignoredVolume, not correctness. 61 correct items is still 61 itemsCut what surfaces per document, and rank what remains
Precision of the flag stream, what it usually means, and the repair it points at

A playbook of ideal sentences flags everything that is not that sentence

Most playbooks begin life as a precedent bank: the wording the firm prefers, saved from the deal where it was won. Wire that directly into a comparison and the rule becomes "flag anything that is not this", which is nearly every clause in nearly every agreement, because counterparties do not draft from your precedent.

The repair is to write each entry as a range with a boundary. Not "notice must be 30 days" but "30 days preferred, 15 to 60 acceptable, below 15 escalate". The boundary is the flag; everything inside it is silence. Where a firm has no written positions to convert, comparing against what it has actually signed is the alternative standard, weighed in judging a clause against a playbook or against your own past deals.

Getting the ranges out of a partner is an elicitation problem, and it fails the same way client-facing questions fail: ask "what is your position on limitation of liability" and you get an essay. Ask "at what number would you refuse to sign" and you get a boundary. The same discipline is set out in writing the drafting questionnaire a client can answer.

The governing-law nit arrives beside the uncapped indemnity

A flag stream with no severity is a list of 61 equal objects, and a reviewer processes an unordered list by sampling it. This is the failure that survives every accuracy improvement, because each individual flag is defensible: the governing law really is not the firm's preference, and it really does not matter next to an indemnity with no cap in sight.

  • Tier by consequence, not by confidence. What the clause can cost if it goes wrong is a property of the clause type and is stable; a model score moves with wording and tells you nothing about exposure.
  • Cap what surfaces per document. 3 material items and a collapsed list of the rest is read; 61 flat items are not, whatever the underlying quality.
  • Give every tier a named consumer. If nobody would act on a tier-3 flag, it should not be a flag — it can live in a report nobody has to clear.
  • Keep the tier stable across releases. A clause that was tier 1 last month and tier 3 today teaches reviewers that the ordering means nothing.

3 phrasings of one accepted position, counted as 3 deviations

This is usually the largest single bucket in the 50, and the cheapest to remove. A firm accepts a mutual confidentiality obligation surviving 3 years; the counterparty writes it as "for a period of three (3) years from the date of disclosure", or "for 36 months", or as an obligation that survives termination by 3 years in a separate survival clause. Same position, 3 hits.

Fix it in the playbook entry rather than in the model: every entry should carry the accepted variants it has already been shown, so each graded false positive permanently removes a class of future ones. That turns the 50-flag exercise into a ratchet instead of a one-off audit, and it is the mechanism that makes a weekly review cycle worth running.

Clause types where the firm has no position, and the honest way to say so

Some entries exist because the taxonomy demanded one, not because anybody has a view. Force majeure in a low-value services agreement is the classic: the entry says "review", the system dutifully flags it every time, and no reviewer has ever changed a word of it. A playbook entry with no position is not a cautious entry — it is an unanswered question that has been converted into recurring work.

Name those clause types explicitly and set them to no position, so they are extracted and stored but never flagged. It is a smaller list than teams expect, usually 4 to 8 types, and removing it often takes 20 per cent off the flag count in an afternoon. Reviewing the list quarterly is enough; a clause type acquires a position the first time a deal turns on it.

A reviewer who skims 61 flags is not being careless. They are correctly inferring that a list nobody ordered was produced by a system that does not know which item matters.

Rewriting one entry so it stops firing on acceptable language

  1. Pull every flag the entry raised in the last 30 days and sort by clause text. Duplicates cluster immediately, and the clusters are the variants you have been paying for.
  2. Write the boundary as a testable condition. A number, a date range, a named party role or an explicit list — anything a person can check without interpreting the entry.
  3. Attach the accepted variants seen in the sample, so the same wording never fires twice. This is the part teams skip, and it is the part that compounds.
  4. Set the severity tier deliberately, using the consequence of the clause rather than how confident the detection was.
  5. Re-run against the graded 50 and compare flag-for-flag. An entry that removes 9 false positives and 1 true one has not improved; it has traded, and the trade needs a lawyer's sign-off.
  6. Version the entry with a date and an owner, so a rule that starts misbehaving can be traced to the change that caused it.

What a quieter flag stream still leaves undone

Precision fixes the reading problem, not the output problem. A well-ranked set of 4 flags still has to become something a counterparty can act on, and where the edit is written into the document body rather than as attributable revisions the negotiation record is destroyed — the separate defect in the suggested edit that lands as clean text.

Nor does calibration survive on its own. Playbook entries drift as deals are done, and the ratchet only works if the graded sample runs on a schedule with a named owner. Wiring that loop — sample, grade, amend the entry, re-run — is ordinary AI automation and agent engineering, and it belongs with the rest of contract review, clause risk and redlining in the work we do with legal teams.

Frequently asked questions

Short answers to the follow-ups this page tends to raise.

Why does an AI contract review tool flag almost every clause?

Because the playbook it compares against is usually written as preferred wording rather than as acceptable ranges. Anything that is not the preferred sentence becomes a deviation, and counterparties never draft from your precedent, so the flag count tracks document length rather than risk. The tell is that flag volume rises with page count and clusters on 2 or 3 clause types.

How many flags should I sample to know whether the problem is calibration or ranking?

50 consecutive flags from live reviews, graded by 1 lawyer into would-have-changed and would-have-accepted. That is enough to separate the two failure modes, which is all the sample has to do. A low would-have-changed share means the flags are wrong and the playbook entries need rewriting; a high share with reviewers still ignoring the panel means the flags are right and unranked.

Is severity tiering better than raising the confidence threshold?

Yes, for this problem. A confidence threshold suppresses items the model was unsure about, which is unrelated to how much a clause can cost you — it will hide an uncertain uncapped indemnity while showing a certain governing-law preference. Tiering by consequence keeps the material items visible regardless of how confidently they were detected, and it stays stable when the model or the prompt changes.

What should a playbook do with clause types the firm has no view on?

Mark them explicitly as no position, so they are extracted and stored but never flagged. An entry that says "review" without stating what would be unacceptable generates work forever and resolves nothing. The list is usually 4 to 8 clause types, and removing them frequently takes a fifth off the flag count without changing a single detection rule.

  • playbook
  • alert fatigue
  • contract review
  • calibration
// shipped work

The work behind this page

Builds from our portfolio that this page draws on.

Read next

Working on something in this space?

Tell us where you are in a sentence or two. We'll tell you honestly whether we're the right team, and what a sensible first slice of the work looks like.

Start the conversation