Extraction is clean on the issued PDFs and useless on the scanned ones
In short
Extraction that works on issued PDFs and fails on scans is a capture problem, not a model problem. The smallest annotation was scanned below the resolution it needs, the sheet is skewed, or the linework has too little contrast against a faded background. Measure those 3 on one sample sheet before tuning anything, because no amount of processing recovers detail that was never captured.
Key takeaways
- Resolution is a property of the smallest text on the sheet, not of the scanner setting.
- At 300 dpi one pixel is 0.085 mm, so 2.5 mm text is about 30 pixels tall. At 150 dpi it is 15.
- A 0.5 degree skew displaces the far edge of an A1 sheet by roughly 7 mm — more than a schedule row.
- An issued PDF has a text layer and a scan does not. Test which one you have before blaming the model.
- Re-scanning is cheaper than processing whenever the original still exists and the sheet will be read repeatedly.
An issued PDF plotted from a CAD or BIM tool carries text as text and lines as vectors, so extraction is a lookup. A scan carries a photograph of those things, and everything downstream depends on how well the photograph was taken. When extraction is clean on one and useless on the other, the difference is nearly always capture: the smallest annotation was scanned below the resolution it needs, the sheet is skewed or warped, or the linework has too little contrast against the paper. All three are measurable on a single sample sheet in about 15 minutes.
Establish which file you actually have before anything else. Select text in the viewer, or run a text extractor and see whether anything comes back. A file that returns nothing is a raster image regardless of the fact that it is a PDF, and a file that returns text on some sheets and nothing on others is a mixed set — the common case for a project reissued over several years.
Measure the sheet before you tune the model
- Take one representative sheet, not the worst one. A plan with a schedule, dimension strings and revision clouds tells you more than a title sheet.
- Read the pixel dimensions and divide by the sheet size in inches. An A1 sheet is 594 by 841 mm, so 7,000 pixels on the long edge is about 210 dpi, whatever the file's metadata claims.
- Zoom to the smallest text on the sheet — usually dimension or note text rather than room names — and measure its height in pixels. That number, not the sheet dpi, is what decides whether it reads.
- Draw a line along a border or a schedule rule and measure its angle. Anything above about 0.3 degrees will start to hurt table extraction.
- Sample the grey level of the linework and of the background nearby. Faded diazo prints often land within 30 levels of each other, which is where binarisation starts making arbitrary choices.
- Look for a screen pattern in solid fills at high zoom. Regular dot structure means the original was halftoned and the scan may be carrying moiré rather than content.
Write those 5 numbers down per sheet batch. They are what tell you whether the set is one problem or four, and whether the answer is a processing pipeline or a scanner.
Resolution is a property of the smallest text, not the sheet
The Tesseract documentation states plainly that it works best on images of at least 300 dpi and suggests rescaling below that. That figure is about text on a page of ordinary letter size. A drawing is a much larger sheet with much smaller text on it, so the sheet-level number can be met while the text is still unreadable — and conversely, a lower sheet dpi can be fine if the annotations are large.
The trap is scanning a reduced print. A full-size sheet reprinted at half size and then scanned at 300 dpi gives you an effective 150 dpi against the original geometry, and no metadata anywhere records that the reduction happened. If a set reads badly and the numbers look fine, ask what was physically on the scanner bed.
Half a degree is enough to break a schedule
Skew is the failure people underestimate, because half a degree is invisible to the eye across a sheet. Across the 841 mm width of an A1 sheet, 0.5 degrees displaces the far edge by about 7 mm. Schedule rows are commonly 4 to 6 mm apart, so by the right-hand edge of the table the text has drifted more than a full row out of line with its own row rule. The Tesseract documentation is explicit that line segmentation quality drops significantly on a skewed page, and a table parser has an even harder time, because it is grouping by horizontal band.
Rolled originals add a second problem that deskewing does not fix: the sheet is not flat, so the angle changes across the page. A global rotation corrects the middle and makes the edges worse. Sheets that have lived in a tube want flattening under weight before scanning, or a scanner with a proper feed rather than a flatbed and a hand. When a schedule survives all this and still comes back wrong, the cause is structural rather than optical, and that is a different diagnosis — the extracted schedule that is one column out of alignment covers it.
Faded prints, negatives and screened fills
- Faded diazo and blueline prints. The line and the background are both mid-tones, so a global threshold either drops the thinnest lines or fills the sheet with speckle. Adaptive thresholding — the Sauvola and adaptive-Otsu style methods that current OCR toolchains offer — is the right tool, because it decides per region rather than per sheet.
- Blueprint negatives. White lines on a dark ground read as noise to a pipeline expecting dark ink on white. Invert first, then threshold. A pipeline that never considers inversion will report a blank sheet with total confidence.
- Halftoned or screened fills. Poché, hatching and shaded zones printed as dot screens interfere with the scan grid and produce moiré that looks like content. Scanning at a higher resolution and descreening beats trying to filter it afterwards.
- Stains, folds and tape. Local damage is local: mask the region, extract the rest, and record that the region was unreadable rather than letting the model guess. A confident wrong value on a fire rating is worse than a blank.
Photographed sheets are their own category and worth separating out, because the failure is geometric rather than tonal. A phone photo of a drawing on a table has perspective distortion, uneven lighting and a focal plane that is sharp in the middle and soft at the corners. Perspective correction against the sheet border recovers a usable image for reading a note or a tag, but not for measurement, and not usually for a dense schedule. Where a photo is the right capture method in the first place — proving a condition existed rather than reading a document — is the argument in photo markup or tablet redline for as-builts.
The point where re-scanning beats any amount of processing
Processing has a ceiling: it cannot add detail that was never captured. If the smallest text is 8 pixels tall, no model, no upscaling and no threshold recovers the digits reliably, and every hour spent tuning is an hour spent proving that. The decision is not technical so much as arithmetic.
| Situation | Decision |
|---|---|
| Original or a full-size print still exists, and the sheet will be queried repeatedly | Re-scan at 400 dpi or better, greyscale, deskewed at capture |
| Only a reduced print survives, text under about 12 pixels | Re-scan will not help. Route to human read, and record what was unreadable |
| Set reads acceptably except for schedules and small annotations | Re-scan the affected sheets only — typically a small fraction of the set |
| Archive material nobody has opened in years | Leave it. Scan on demand, at quality, when a sheet is actually needed |
That last row is the one teams get wrong most often, and it is the same shape of judgement as when tagging costs more than the tools it protects: the cost of capture is certain and immediate, the benefit is probabilistic and later. Scanning 900 archive sheets to read 40 of them is a bad trade, and it is a trade people make because the set feels like the unit of work. The sheet is the unit of work.
Processing cannot add detail that was never captured. Past a certain pixel height you are not tuning a model, you are proving it.
What a clean scan still will not give you
A well-captured sheet moves you from an image problem to an ordinary document-AI problem, and those have their own limits. Title blocks and schedule tables read well; dense symbol fields and handwritten markup do not, and the ranking of what is dependable is in which parts of a drawing set machines read reliably. Specification books scan far better than drawings, being ordinary text at ordinary sizes, which is why turning a spec book into a submittal list is tractable on an old project where the drawings are not.
Two practical rules for the pipeline. Record a per-sheet confidence and the measured inputs — dpi, smallest text height, skew, contrast — alongside every extraction, so a downstream reader can tell a value that was read from a value that was guessed. And keep the extraction attached to what happens next: a register built from an old set is only useful if its status is maintained in one place, which is the failure behind a submittal log that says approved while the inbox says otherwise.
If a vendor tells you their extraction handles any drawing, ask what it does with a 150 dpi blueline scan of a reduced print — the answer separates the tools that measure their input from the ones that do not, which is one of the questions in choosing an AI development partner. Pipelines of this kind are what we scope under AI agents and automation, inside drawings, specs and construction document intelligence, part of our construction and contracting work.
Frequently asked questions
Short answers to the follow-ups this page tends to raise.
What resolution should construction drawings be scanned at for OCR?
Measure the smallest text rather than picking a sheet-level number. Aim for at least 25 to 30 pixels of cap height on the smallest annotation you need to read, which on typical 2.5 mm drawing text means 300 dpi and on 1.8 mm dimension text means 400 dpi or more. Scanning a reduced print halves your effective resolution without anything recording that it happened.
Why does OCR work on some sheets of a set and not others?
Usually because the set is mixed. Sheets reissued from CAD carry a real text layer and extract perfectly, while sheets that came back as scans carry only pixels. Run a text extractor across every sheet and split the set into those that return text and those that return nothing — the two halves need completely different handling, and treating them as one set is what makes results look random.
Can a model read a photograph of a drawing taken on a phone?
For a note, a tag or a title block, often yes, after perspective correction against the sheet border. For a dense schedule or anything dimensional, usually not: a phone photo carries perspective distortion, uneven lighting and corner softness that no correction fully removes. Treat photos as evidence that a condition existed, not as a substitute for a scan of the document.
Is it worth re-scanning old drawings, or should we process what we have?
Re-scan when the original or a full-size print still exists and the sheet will be read more than once — it is the cheapest fix available, and it removes the problem instead of managing it. Do not re-scan a whole archive on principle: scan the sheets that are actually queried, at quality, and leave the rest until somebody needs them.
- document ai
- scanning
- as-builts
- extraction
The work behind this page
Builds from our portfolio that this page draws on.
GroundUp
A construction project-management command centre for general contractors that keeps schedule, RFIs, budget and the field log in one place — and maps the critical-path recovery the moment a job slips.
Real EstateAskVault
An AI internal knowledge-search platform that answers employee questions from your own docs — grounded in citations, with knowledge gaps surfaced and deflection tracked.
Productivity AIRead next
- The extracted equipment schedule is one column out of alignmentValidate three known rows against the sheet by tag number rather than by position, then run the four post-extraction rules that catch a shift before anyone prices the schedule.diagnostic
- The drawing comparison marks every sheet as changedWhen a comparison returns differences on every sheet, it is usually telling you it never managed to line the two files up. Three sheets you know did not change will settle it in about ten minutes.diagnostic
- The title block is the metadata layer of a sheet setEvery document system built over a sheet set is really built over the title block. Some of its fields are dependable across consultants and some are not.definition
- Deferred submittals: the items the register has to hold openA deferred item is an obligation with a later trigger and a reviewer outside your contract. A register that files it under 'not started' has already lost it.definition
- Order of precedence when the drawing and the specification conflictPrecedence is a contract term, not an industry constant. Until somebody reads the clause on this project, no conflict-detection tool can rank what it finds.definition
- The submittal log says approved and the inbox says otherwiseRun an ageing report by ball-in-court and put the last dated communication beside every open row. The rows where those two disagree are the drift, and the pattern names the cause.diagnostic
Working on something in this space?
Tell us where you are in a sentence or two. We'll tell you honestly whether we're the right team, and what a sensible first slice of the work looks like.
Start the conversation