Every clause comes back flagged, so the reviewer stops reading the flags
In short
When a contract review tool flags too many clauses, the model is almost never at fault. A playbook written as ideal sentences flags every acceptable variant of them, and one with no severity tier puts a governing-law nit beside an uncapped indemnity. Grade 50 flags would-have-changed or would-have-accepted; the split names the repair.
Key takeaways
- Grade 50 flags before touching anything. Wrong flags and unranked flags look identical and need opposite repairs.
- A playbook written as preferred wording flags every acceptable paraphrase of that wording, which is most of the market.
- With no severity tier, a governing-law nit and an uncapped indemnity arrive as the same object, so both get skimmed.
- Semantic equivalence is the largest single bucket: 3 phrasings of one accepted position counted as 3 deviations.
- Clause types where the firm has no position should say so. Silence dressed as a flag is the worst output available.
- Precision above roughly 70 per cent turns the problem from calibration into ranking, and the fix changes accordingly.
A reviewer who opens a 40-page agreement, sees 61 flags and closes the panel has made a rational decision. Reading 61 items to find the 4 that matter is slower than reading the contract, and everyone works this out inside a week. The tool is not being ignored because lawyers distrust software. It is being ignored because it has stopped carrying information.
Almost all of this is playbook calibration, and almost none of it is the model. The playbook is a set of written positions; if those positions are expressed as exact sentences rather than acceptable ranges, every reasonable paraphrase in the market becomes a deviation and the flag count tracks document length instead of risk.
50 flags, 1 lawyer, and 2 columns
The measurement comes first because the two failure modes are indistinguishable from the flag count alone. A tool producing 61 useless flags and a tool producing 61 correct flags in no useful order both look like noise to the reviewer, and they need opposite repairs.
- Take 50 consecutive flags from real reviews, not from a demo set. Consecutive matters: cherry-picking by clause type reproduces whatever bias created the problem.
- Have 1 lawyer mark each flag would-have-changed or would-have-accepted, answering only whether they would have negotiated this wording, with no third option. A maybe column collapses the measurement.
- Discard flags raised against text that was not operative before scoring anything. A flag on a superseded clause is a different defect entirely, diagnosed in the review that is right about a superseded version, and counting it as a false positive sends you to the wrong fix.
- Record the clause type against every flag. The distribution matters more than the total: 34 of 50 landing on 3 clause types is a playbook-entry problem, while an even spread across 15 types is a threshold problem.
- Compute the would-have-changed share. That single number is the precision of the flag stream, and it routes the rest of this page.
- Ask the same lawyer for the flags they expected and did not get. Misses never appear in a false-positive count, and a playbook tuned only on precision quietly stops catching things.
What the would-have-changed share tells you to do next
The bands below are working thresholds rather than published standards, and they are useful because each one implies a different piece of work. Set your own after 2 rounds of measurement; what matters is that the number decides the action rather than the other way round.
| Would-have-changed share | What it usually means | Where to work first |
|---|---|---|
| Below 30 per cent | The playbook encodes preferred wording, not acceptable ranges | Rewrite the 5 highest-volume entries as ranges before anything else |
| 30 to 60 per cent | Real deviations exist but sit under a pile of paraphrase hits | Semantic equivalence: collapse variants of positions you already accept |
| 60 to 80 per cent | Flags are broadly right and arrive in no defensible order | Add severity tiers; stop changing the detection at all |
| Above 80 per cent, still ignored | Volume, not correctness. 61 correct items is still 61 items | Cut what surfaces per document, and rank what remains |
A playbook of ideal sentences flags everything that is not that sentence
Most playbooks begin life as a precedent bank: the wording the firm prefers, saved from the deal where it was won. Wire that directly into a comparison and the rule becomes "flag anything that is not this", which is nearly every clause in nearly every agreement, because counterparties do not draft from your precedent.
The repair is to write each entry as a range with a boundary. Not "notice must be 30 days" but "30 days preferred, 15 to 60 acceptable, below 15 escalate". The boundary is the flag; everything inside it is silence. Where a firm has no written positions to convert, comparing against what it has actually signed is the alternative standard, weighed in judging a clause against a playbook or against your own past deals.
Getting the ranges out of a partner is an elicitation problem, and it fails the same way client-facing questions fail: ask "what is your position on limitation of liability" and you get an essay. Ask "at what number would you refuse to sign" and you get a boundary. The same discipline is set out in writing the drafting questionnaire a client can answer.
The governing-law nit arrives beside the uncapped indemnity
A flag stream with no severity is a list of 61 equal objects, and a reviewer processes an unordered list by sampling it. This is the failure that survives every accuracy improvement, because each individual flag is defensible: the governing law really is not the firm's preference, and it really does not matter next to an indemnity with no cap in sight.
- Tier by consequence, not by confidence. What the clause can cost if it goes wrong is a property of the clause type and is stable; a model score moves with wording and tells you nothing about exposure.
- Cap what surfaces per document. 3 material items and a collapsed list of the rest is read; 61 flat items are not, whatever the underlying quality.
- Give every tier a named consumer. If nobody would act on a tier-3 flag, it should not be a flag — it can live in a report nobody has to clear.
- Keep the tier stable across releases. A clause that was tier 1 last month and tier 3 today teaches reviewers that the ordering means nothing.
3 phrasings of one accepted position, counted as 3 deviations
This is usually the largest single bucket in the 50, and the cheapest to remove. A firm accepts a mutual confidentiality obligation surviving 3 years; the counterparty writes it as "for a period of three (3) years from the date of disclosure", or "for 36 months", or as an obligation that survives termination by 3 years in a separate survival clause. Same position, 3 hits.
Fix it in the playbook entry rather than in the model: every entry should carry the accepted variants it has already been shown, so each graded false positive permanently removes a class of future ones. That turns the 50-flag exercise into a ratchet instead of a one-off audit, and it is the mechanism that makes a weekly review cycle worth running.
Clause types where the firm has no position, and the honest way to say so
Some entries exist because the taxonomy demanded one, not because anybody has a view. Force majeure in a low-value services agreement is the classic: the entry says "review", the system dutifully flags it every time, and no reviewer has ever changed a word of it. A playbook entry with no position is not a cautious entry — it is an unanswered question that has been converted into recurring work.
Name those clause types explicitly and set them to no position, so they are extracted and stored but never flagged. It is a smaller list than teams expect, usually 4 to 8 types, and removing it often takes 20 per cent off the flag count in an afternoon. Reviewing the list quarterly is enough; a clause type acquires a position the first time a deal turns on it.
A reviewer who skims 61 flags is not being careless. They are correctly inferring that a list nobody ordered was produced by a system that does not know which item matters.
Rewriting one entry so it stops firing on acceptable language
- Pull every flag the entry raised in the last 30 days and sort by clause text. Duplicates cluster immediately, and the clusters are the variants you have been paying for.
- Write the boundary as a testable condition. A number, a date range, a named party role or an explicit list — anything a person can check without interpreting the entry.
- Attach the accepted variants seen in the sample, so the same wording never fires twice. This is the part teams skip, and it is the part that compounds.
- Set the severity tier deliberately, using the consequence of the clause rather than how confident the detection was.
- Re-run against the graded 50 and compare flag-for-flag. An entry that removes 9 false positives and 1 true one has not improved; it has traded, and the trade needs a lawyer's sign-off.
- Version the entry with a date and an owner, so a rule that starts misbehaving can be traced to the change that caused it.
What a quieter flag stream still leaves undone
Precision fixes the reading problem, not the output problem. A well-ranked set of 4 flags still has to become something a counterparty can act on, and where the edit is written into the document body rather than as attributable revisions the negotiation record is destroyed — the separate defect in the suggested edit that lands as clean text.
Nor does calibration survive on its own. Playbook entries drift as deals are done, and the ratchet only works if the graded sample runs on a schedule with a named owner. Wiring that loop — sample, grade, amend the entry, re-run — is ordinary AI automation and agent engineering, and it belongs with the rest of contract review, clause risk and redlining in the work we do with legal teams.
Frequently asked questions
Short answers to the follow-ups this page tends to raise.
Why does an AI contract review tool flag almost every clause?
Because the playbook it compares against is usually written as preferred wording rather than as acceptable ranges. Anything that is not the preferred sentence becomes a deviation, and counterparties never draft from your precedent, so the flag count tracks document length rather than risk. The tell is that flag volume rises with page count and clusters on 2 or 3 clause types.
How many flags should I sample to know whether the problem is calibration or ranking?
50 consecutive flags from live reviews, graded by 1 lawyer into would-have-changed and would-have-accepted. That is enough to separate the two failure modes, which is all the sample has to do. A low would-have-changed share means the flags are wrong and the playbook entries need rewriting; a high share with reviewers still ignoring the panel means the flags are right and unranked.
Is severity tiering better than raising the confidence threshold?
Yes, for this problem. A confidence threshold suppresses items the model was unsure about, which is unrelated to how much a clause can cost you — it will hide an uncertain uncapped indemnity while showing a certain governing-law preference. Tiering by consequence keeps the material items visible regardless of how confidently they were detected, and it stays stable when the model or the prompt changes.
What should a playbook do with clause types the firm has no view on?
Mark them explicitly as no position, so they are extracted and stored but never flagged. An entry that says "review" without stating what would be unacceptable generates work forever and resolves nothing. The list is usually 4 to 8 clause types, and removing them frequently takes a fifth off the flag count without changing a single detection rule.
- playbook
- alert fatigue
- contract review
- calibration
The work behind this page
Builds from our portfolio that this page draws on.
Brief Forge
Contract review AI for solo lawyers and small firms — extract, score, and redline contracts in minutes.
Legal TechNotewell
An AI meeting assistant that records and transcribes every meeting, extracts the decisions and action items, assigns owners and due dates, and tracks follow-through until it's done.
Productivity AIRead next
- Fallback positions: preferred, acceptable and walk-away, in one ladderA playbook that records only your ideal clause can say the counterparty deviated. It cannot say what to send back, which is the part a negotiator needs.definition
- The extraction found the indemnity and got the direction wrongThe clause was found. The indemnifying party, the triggering events and the defence mechanics were not, and a record reading "indemnity: present" is worse than an empty field.diagnostic
- Defined terms: no clause can be judged without its definitionA capitalised term carries whatever meaning the contract assigns it, wherever that assignment happens to live. Judge the clause without resolving it and you have judged a sentence you have not read.definition
- Extraction works on the clean draft and fails on the signed executed copyThe pipeline scores well on the draft and loses clause boundaries on the file that was actually signed. Six pathologies that exist only on executed documents, and what the system should refuse to answer.diagnostic
- Order of precedence: the clause that decides which document winsA precedence clause ranks a stack of documents. It is also why a single file is an incomplete unit of review: the same words bind or not, depending on which instrument controls.definition
- The liability cap is three fields, and only one of them is a numberThe cap is not a number. It is an amount, a basis that computes the amount, and a list of obligations that sit outside both — and the third one decides how bad the worst day gets.definition
Working on something in this space?
Tell us where you are in a sentence or two. We'll tell you honestly whether we're the right team, and what a sensible first slice of the work looks like.
Start the conversation