Real Estate & PropTech// definition

Confidence on an extracted lease field: what the number measures

In short

A confidence score on an extracted lease field measures how consistently the system reached the value — agreement between passes, retrieval strength, layout certainty — and never whether the value is true. Confidently wrong is the ordinary failure, and the threshold you pick states how many fields a person can check in a day.

Key takeaways

  • Confidence is a statement about the extraction process, not about the lease. Nothing in it reaches the document's truth.
  • A clean, well-laid-out table scores high whether or not the schedule it holds is still the operative one.
  • Thresholds are set by capacity: 40 fields an hour across 300 leases decides the cut-off, not a quality target.
  • A score may queue work. It may not authorise a payment, a notice, a signature or a number in an investor report.

A confidence score on an extracted lease field tells you how consistently the system reached that value. It does not tell you whether the value is right. Those are different questions, and the gap between them is where trust in an abstraction pipeline is built or quietly lost.

The useful consequence is narrow: the number orders a review queue well and authorises nothing. A field at 0.93 and one at 0.61 differ in how likely they are to reward attention, not in whether you may pay against them.

The 3 signals usually hiding behind one number

Vendors and in-house pipelines alike expose one figure per field, normally blended from the 3 signals below. Ask which yours uses: the answer changes what a low score is evidence of.

SignalWhat it measuresWhat a low value meansWhat a high value does not prove
Agreement between passesWhether 2 or more independent reads produced the same valueThe document is ambiguous, or the field is genuinely stated twiceThat the agreed value came from the right clause
Retrieval strengthHow well the located passage matched what the field asked forThe language is absent, or sits in an exhibit nobody searchedThat the passage is operative rather than a superseded recital
Layout and character certaintyHow cleanly the page transcribed — skew, stamps, handwriting, merged cellsA scan, a fax generation, or a table whose structure collapsedAnything about meaning. A crisp scan of the wrong page scores well
The signals a per-field confidence figure is typically built from

None of the 3 has a channel through which the truth of the lease could reach it — the distinction drawn for model outputs generally in what a confidence score actually measures. Leases add layout, and layout is what gets mistaken for comprehension.

The high-confidence rent that came from the wrong document

The expensive failure is a high score on a value read cleanly from the wrong place. A rent schedule in a crisp table on page 4 of the original lease scores high for as long as the file exists, including the 3 years after an amendment replaced it — the document-set problem in the abstract showing the original rent, not the amended one.

The second version is a compound field pretending to be simple. An expense cap extracted as 5 per cent is confidently correct and useless without its basis: a cap accumulating unused headroom behaves nothing like one that does not, per cumulative and non-cumulative expense caps. The number was never the field.

The threshold states how many fields a person can check

Teams set thresholds as though picking a quality bar. The constraint is arithmetic. At 40 fields an hour, a 300-lease backlog of 40 fields each is 12,000 fields and 300 reviewer-hours. The threshold is the cut-off that brings that queue inside the hours you have.

  1. Measure the review rate first. Fields verified per hour, on your documents, by whoever will really do it.
  2. Rank fields by consequence rather than by score. Rent, dates, option windows and cap basis outrank a use clause.
  3. Set the cut-off so the queue fits capacity, and treat it as provisional. It moves when the backlog clears or the document mix changes.
  4. Sample above the line as well as below it. Checking some of what the threshold passed is the only way to learn what it is letting through.

How that queue is staffed and what its exit criteria are belongs to the review queue that makes extraction safe to rely on. Whether you have a per-field score at all depends on owning the pipeline rather than buying a finished abstract from a desk — the trade in an outsourced abstraction desk against your own pipeline.

4 decisions the score is not allowed to make

  • Paying. A rent, a true-up or a percentage-rent figure is paid against a document and a citation somebody opened.
  • Notifying. Exercising or declining an option is irreversible on a date, so a named person verifies the field or no notice fires.
  • Signing. Anything becoming a representation — an estoppel response, a lender certification — leaves the pipeline entirely.
  • Reporting. A portfolio number reaching an investor carries the assumption that a person stands behind it, and an average is not that person.

The same caution applies to a headline accuracy figure, since an average across 40 fields of unequal consequence hides which were wrong: reading accuracy claims in lease abstraction. Designing the field set, the score and the queue as one system is MVP and product engineering, within lease abstraction and property document AI for real estate teams.

Frequently asked questions

Short answers to the follow-ups this page tends to raise.

Does a high confidence score mean an extracted lease field is correct?

No. It means the system reached that value consistently — 2 passes agreed, the passage matched, the page transcribed cleanly. A crisp table on the wrong page, or the original rent schedule in a lease later amended, scores high indefinitely. Correctness is a property of the document set; confidence is a property of the process.

What confidence threshold should send a lease field to a reviewer?

Whichever one makes the queue fit your review capacity. Measure how many fields a reviewer verifies per hour against citations, multiply by fields per lease and leases in the backlog, then set the cut-off so the queue can be cleared. A vendor's default describes their document mix and staffing, not yours.

Should every field use the same confidence threshold?

No, because the fields carry unequal consequences. Rent, commencement and expiry dates, option windows and the basis of an expense cap control money or rights, and deserve review at a level a use clause does not. One global threshold spends identical effort on the field nobody acts on and the one that sets the payment.

Can a confidence score be used to auto-approve extracted data?

Only where being wrong is cheap and reversible. Payment, notice, signature and external reporting each need a person who opened the citation, whatever the number says. Auto-approval is defensible for descriptive fields with no downstream consumer — and such a field usually should not have been in the schema.

  • confidence
  • lease abstraction
  • extraction
  • review routing
// shipped work

The work behind this page

Builds from our portfolio that this page draws on.

Read next

Working on something in this space?

Tell us where you are in a sentence or two. We'll tell you honestly whether we're the right team, and what a sensible first slice of the work looks like.

Start the conversation