Confidence on an extracted lease field: what the number measures
In short
A confidence score on an extracted lease field measures how consistently the system reached the value — agreement between passes, retrieval strength, layout certainty — and never whether the value is true. Confidently wrong is the ordinary failure, and the threshold you pick states how many fields a person can check in a day.
Key takeaways
- Confidence is a statement about the extraction process, not about the lease. Nothing in it reaches the document's truth.
- A clean, well-laid-out table scores high whether or not the schedule it holds is still the operative one.
- Thresholds are set by capacity: 40 fields an hour across 300 leases decides the cut-off, not a quality target.
- A score may queue work. It may not authorise a payment, a notice, a signature or a number in an investor report.
A confidence score on an extracted lease field tells you how consistently the system reached that value. It does not tell you whether the value is right. Those are different questions, and the gap between them is where trust in an abstraction pipeline is built or quietly lost.
The useful consequence is narrow: the number orders a review queue well and authorises nothing. A field at 0.93 and one at 0.61 differ in how likely they are to reward attention, not in whether you may pay against them.
The 3 signals usually hiding behind one number
Vendors and in-house pipelines alike expose one figure per field, normally blended from the 3 signals below. Ask which yours uses: the answer changes what a low score is evidence of.
| Signal | What it measures | What a low value means | What a high value does not prove |
|---|---|---|---|
| Agreement between passes | Whether 2 or more independent reads produced the same value | The document is ambiguous, or the field is genuinely stated twice | That the agreed value came from the right clause |
| Retrieval strength | How well the located passage matched what the field asked for | The language is absent, or sits in an exhibit nobody searched | That the passage is operative rather than a superseded recital |
| Layout and character certainty | How cleanly the page transcribed — skew, stamps, handwriting, merged cells | A scan, a fax generation, or a table whose structure collapsed | Anything about meaning. A crisp scan of the wrong page scores well |
None of the 3 has a channel through which the truth of the lease could reach it — the distinction drawn for model outputs generally in what a confidence score actually measures. Leases add layout, and layout is what gets mistaken for comprehension.
The high-confidence rent that came from the wrong document
The expensive failure is a high score on a value read cleanly from the wrong place. A rent schedule in a crisp table on page 4 of the original lease scores high for as long as the file exists, including the 3 years after an amendment replaced it — the document-set problem in the abstract showing the original rent, not the amended one.
The second version is a compound field pretending to be simple. An expense cap extracted as 5 per cent is confidently correct and useless without its basis: a cap accumulating unused headroom behaves nothing like one that does not, per cumulative and non-cumulative expense caps. The number was never the field.
The threshold states how many fields a person can check
Teams set thresholds as though picking a quality bar. The constraint is arithmetic. At 40 fields an hour, a 300-lease backlog of 40 fields each is 12,000 fields and 300 reviewer-hours. The threshold is the cut-off that brings that queue inside the hours you have.
- Measure the review rate first. Fields verified per hour, on your documents, by whoever will really do it.
- Rank fields by consequence rather than by score. Rent, dates, option windows and cap basis outrank a use clause.
- Set the cut-off so the queue fits capacity, and treat it as provisional. It moves when the backlog clears or the document mix changes.
- Sample above the line as well as below it. Checking some of what the threshold passed is the only way to learn what it is letting through.
How that queue is staffed and what its exit criteria are belongs to the review queue that makes extraction safe to rely on. Whether you have a per-field score at all depends on owning the pipeline rather than buying a finished abstract from a desk — the trade in an outsourced abstraction desk against your own pipeline.
4 decisions the score is not allowed to make
- Paying. A rent, a true-up or a percentage-rent figure is paid against a document and a citation somebody opened.
- Notifying. Exercising or declining an option is irreversible on a date, so a named person verifies the field or no notice fires.
- Signing. Anything becoming a representation — an estoppel response, a lender certification — leaves the pipeline entirely.
- Reporting. A portfolio number reaching an investor carries the assumption that a person stands behind it, and an average is not that person.
The same caution applies to a headline accuracy figure, since an average across 40 fields of unequal consequence hides which were wrong: reading accuracy claims in lease abstraction. Designing the field set, the score and the queue as one system is MVP and product engineering, within lease abstraction and property document AI for real estate teams.
Frequently asked questions
Short answers to the follow-ups this page tends to raise.
Does a high confidence score mean an extracted lease field is correct?
No. It means the system reached that value consistently — 2 passes agreed, the passage matched, the page transcribed cleanly. A crisp table on the wrong page, or the original rent schedule in a lease later amended, scores high indefinitely. Correctness is a property of the document set; confidence is a property of the process.
What confidence threshold should send a lease field to a reviewer?
Whichever one makes the queue fit your review capacity. Measure how many fields a reviewer verifies per hour against citations, multiply by fields per lease and leases in the backlog, then set the cut-off so the queue can be cleared. A vendor's default describes their document mix and staffing, not yours.
Should every field use the same confidence threshold?
No, because the fields carry unequal consequences. Rent, commencement and expiry dates, option windows and the basis of an expense cap control money or rights, and deserve review at a level a use clause does not. One global threshold spends identical effort on the field nobody acts on and the one that sets the payment.
Can a confidence score be used to auto-approve extracted data?
Only where being wrong is cheap and reversible. Payment, notice, signature and external reporting each need a person who opened the citation, whatever the number says. Auto-approval is defensible for descriptive fields with no downstream consumer — and such a field usually should not have been in the schema.
- confidence
- lease abstraction
- extraction
- review routing
The work behind this page
Builds from our portfolio that this page draws on.
AI Lease Management
AI-powered commercial real estate lease management for multi-brand operators — automates lease data extraction, obligation tracking, and portfolio intelligence.
Real EstateScanQueue
An AI radiology worklist that flags suspected critical findings on incoming CT, MR and X-ray studies and orders every read by acuity and SLA — so the sickest patient is read first, not FIFO.
Healthcare AIRead next
- The abstract shows the original rent, not the rent being paidA billed rent and an abstracted rent that differ by one clean step point at the document set, not the model. One count separates a missing amendment from a merge that ran in the wrong order.diagnostic
- The extracted clause cites page 34, and page 34 says nothing of the kindBroken provenance is its own defect class: a wrong value spoils one field, while a citation that will not resolve makes every field on the record unverifiable.diagnostic
- Co-tenancy: the clause that can switch fixed rent offCo-tenancy makes one tenant's rent depend on other tenants trading. It is the clearest case where a lease field is not a number but a small state machine attached to the rent line.definition
- Execution, delivery, rent commencement: 3 start dates, not one fieldExecution binds the parties, delivery usually starts the term, rent commencement starts the money. Collapse them into a single start date and at least one downstream clock is wrong.definition
- Extraction returns no rent schedule, because the schedule is in Exhibit BAn empty rent schedule almost never means a lease without rent steps. Compare the lease's own exhibit list against the exhibits in the file, and structure failure separates from genuine absence in a minute.diagnostic
- Lease abstract: the summary someone acts on without opening the leaseAn abstract is not a shorter lease. It is the record an operations team runs a portfolio from, and the test it has to pass is whether the week's decisions can be made from it alone.definition
Working on something in this space?
Tell us where you are in a sentence or two. We'll tell you honestly whether we're the right team, and what a sensible first slice of the work looks like.
Start the conversation