The automated count finds devices on some sheets and none on others
In short
When symbol recognition works on some sheets and returns nothing on others, the variable is the sheet: vector or scanned, which consultant drew the symbol, and what the symbol overlaps. Build one row per sheet recording origin, whether text is selectable, detected count and a hand count of one sample area. The empty sheets then sort themselves into three groups.
Key takeaways
- Diagnose per sheet, never per set. A set-level accuracy number hides the sheets that returned nothing at all.
- Selectable text is the fastest substrate test there is: no selectable text means a raster sheet and a different pipeline.
- At 150 dpi a 3 mm symbol is about 18 pixels across, which is why scanned sheets lose fill and hatch distinctions first.
- A count that comes back high is usually one area appearing on both a plan and its enlarged sheet, not a detection error.
- Legend drift between consultants is a per-job symbol map, not a model problem, and it has to be maintained by a person.
- Dedicated takeoff engines are strong here. The work worth buying is sheet triage and the export, not a better detector.
If automated counting returns sensible numbers on most sheets and nothing on a few, the model is not the variable — the sheets are. A drawing set is assembled from several consultants' files, produced on different days by different software, and often part-scanned. Detection is dependable on some of those substrates and impossible on others, so the count fails sheet by sheet rather than degrading smoothly across the job.
This is dangerous because the output looks credible. Totals arrive, found symbols are marked on the sheet, and nothing announces that level 3 came back empty. So the first job is not to fix anything. It is to find out which sheets are wrong and by how much, because the causes need three different responses and two of them are not code changes.
Build the per-sheet detection table before you touch a setting
One row per sheet, filled in for every sheet the condition was supposed to cover — including the ones that returned zero, which are the rows the tool will not volunteer. On an 80-sheet set this is about an hour of work, and it ends the argument about whether the tool is any good.
- List every sheet in scope for the count, from the drawing register rather than from the tool's output. Sheets missing from the output are the finding, so a list built from the output cannot show them.
- Record each sheet's origin: which consultant issued it, and whether it arrived as an issued PDF, a print-to-PDF or a scan. The transmittal usually tells you.
- Try to select a room name or a note with the text cursor. Selectable text means vector geometry underneath; no selectable text means a raster image, and that single test predicts most of the result.
- Record the detected count per sheet as the tool reports it.
- Hand count one sample area per sheet — a single apartment, one 10 m by 10 m grid square, or one room type — and record both the hand count and the detected count for that same area.
- Compute detected divided by hand-counted for the sample. Above about 0.97 the sheet is fine; 0.4 to 0.9 is partial detection; below 0.1 is a sheet the extractor never read at all. Those three bands are three different problems.
- Sort the table by that ratio and look at what the bottom of the list has in common. It is almost always one consultant, one issue date, or one file origin.
Reading the table: pattern to cause
The pattern in the table names the cause faster than any inspection of the model does. Six patterns cover nearly everything.
| Pattern | Most likely cause | Confirming test | What to do |
|---|---|---|---|
| Zero detections, no selectable text | The sheet is a scan or an image-only export | Zoom to 400 percent — raster edges go soft, vector stays crisp | Route to the raster path or hand count; do not re-tune the detector |
| Zero detections, text selects fine | That consultant draws the device differently, or it is a block the library has never seen | Put the symbol from a working sheet beside the symbol from the empty one | Add the variant to the job's symbol map and re-run that sheet |
| Ratio 0.4 to 0.9, evenly spread | Symbols collide with hatching, dimension strings or leaders | Sample one dense area and one sparse area on the same sheet; the gap widens with density | Suppress non-symbol layers if the file is vector, otherwise switch that sheet to assisted counting |
| Plans correct, enlarged plans and details empty | The device is drawn only in a keyed detail, at a different scale | Read the scale note under the detail bubble | Count details as a separate condition with its own scope |
| Count higher than the hand count | The same area appears on both a plan and its enlarged sheet | Look for the match line or the key plan on the enlarged sheet | Set sheet scope on the condition rather than deleting marks by hand |
| One discipline clean, another empty | Legend drift — same device, different symbol between consultants | Open the two legends side by side | Maintain a per-consultant symbol map for the job |
One set, two substrates
The most common finding is that the set is not one thing. The architectural sheets arrive as clean vector exports, the services sheets come from a consultant who scans marked-up prints, and an addendum sheet was photographed on somebody's phone. What the extractor is reading changes completely between those cases, which is the subject of what a quantity extractor is actually reading when it opens your PDF — and the failure mode on scans is covered in extraction that works on issued sheets and fails on scans.
Symbol variants, overlap and legend drift
Where the substrate is fine and the sheet still comes back short, three drawing-side causes account for most of it.
- Consultant-specific variants. A duplex receptacle is drawn one way by one electrical engineer and another way by the next, and both are correct practice. A detector tuned on one job's symbols quietly under-reads the next job's, which is why an accuracy figure quoted without naming the drawing population means very little — see what takeoff accuracy claims are measuring.
- Overlap with linework. Symbols sitting on hatching, inside dimension strings, or under leader text lose the closed outline the detector matches on. The tell is that the miss rate rises with drawing density: sparse rooms count perfectly, plant rooms do not.
- Legend drift across disciplines. The same physical device appears in the electrical legend and the fire legend with different graphics, and only one of the two is in your symbol map. The fix is a per-job map maintained by a person, not a smarter model.
- Symbols that are text. Some consultants tag devices with a text label and no distinct glyph. That is a text-extraction problem wearing the costume of a symbol problem, and it wants a different tool — the ranking in which parts of a drawing set machines read reliably is the map to that.
The devices that were never on the plan
A sheet can be perfectly readable and still under-count, because the device is not drawn there. This is the group that survives every technical fix and lands in the estimate as a missing scope item.
- Riser diagrams and single-line schematics, which carry devices that appear on no floor plan at all.
- Keyed enlarged details, where a typical toilet or a typical patient room is drawn once and referenced from twenty locations.
- Equipment and device schedules, where a table row is the only record that six of something exist.
- Notes carrying multipliers — typical of 6, provide at each column, refer to detail — which are prose instructions to multiply, and no counter reads them as such.
A detector that finds nothing on a sheet is announcing a problem. A detector that finds 80 percent of a sheet is the expensive case, because 80 percent looks like a number you can price.
What to do with a sheet that will not count
- Split the set rather than averaging it. Countable sheets go through automation; the rest are marked for assisted counting, and the split is recorded on the estimate. The rule for drawing that line is in assisted counting versus full automatic takeoff.
- Ask for the vector files before doing anything clever. On a live bid this is a same-day request to the design team and it removes the problem entirely on the sheets it succeeds for.
- Record a qualification for every sheet counted by hand or excluded, in the words the bid form will accept. An unrecorded assumption is what turns a takeoff gap into an argument after award.
- Reconcile the count against a second source — the device schedule, the panel schedule, the spec's product list — before pricing. Two independent sources agreeing beats one source at higher confidence.
- Carry the sheet-level result through to the export, so the quantity that lands on a cost code says how it was produced. Mapping takeoff quantities onto cost codes for export is where that provenance either survives or is lost.
One thing worth saying plainly: dedicated takeoff engines compete hard on symbol recognition, and several are good at it. We do not claim parity with a product whose whole roadmap is counting symbols. What is usually missing around those engines is everything else — the sheet inventory proving which sheets were covered, the exception report that flags a zero-detection sheet before it gets priced, and the mapping into the estimating and accounting systems.
The consequences run downstream in two directions. A count that is short on two sheets becomes a budget line nobody funded, which shows up later as a job budget that does not agree with the estimate that won it. And an unrecorded per-sheet difference is one of the quieter reasons two estimators return different quantities from one set, because each has silently made their own decision about the sheets that would not count.
The durable fix is a scheduled check, not a platform: inventory the sheets, run the count, flag any sheet below your ratio threshold, and email the estimator before work starts. That is the same exception-report shape that stops an expired insurance certificate being found after the sub is on site, and the kind of narrow build we scope under internal tools and ops. The rest of this silo sits under preconstruction, takeoff and estimating, within our construction and contracting work.
Frequently asked questions
Short answers to the follow-ups this page tends to raise.
Why does symbol recognition work on some drawings and not others?
Because the sheets are not the same kind of file. A set typically mixes vector exports, print-to-PDF sheets and scans, and detection quality collapses in that order. Add consultant-specific symbol variants and symbols overlapping hatching, and you get a result that is excellent on half the set and empty on the rest. Test each sheet for selectable text first — it separates the two populations in seconds.
How do I check whether an automated takeoff count is trustworthy?
Hand count one sample area on every sheet and compare it with the tool's count for that same area. A set-level accuracy figure cannot tell you that one floor returned nothing, and that is the failure that costs money. The sample does not need to be large — one apartment, one grid square or one room type per sheet is enough to place the sheet in a band.
Can a scanned drawing be counted automatically at all?
Sometimes, if the scan is clean and high enough resolution, and the symbols are large and well separated. Below roughly 200 dpi at full sheet size a small symbol occupies too few pixels for its fill or hatch to survive, and no processing recovers detail the scan never captured. At that point re-scanning at 300 dpi, or asking the consultant for the vector file, is cheaper than any amount of tuning.
The automated count came back higher than the manual one. What causes that?
Usually the same physical area appearing on more than one sheet — a plan and its enlarged version, or two sheets overlapping at a match line. The detector counts both because it has no idea they are the same room. Fix it by setting sheet scope on the condition rather than deleting marks by hand, because deletions do not survive a re-run after the next addendum.
- takeoff
- symbol recognition
- drawing quality
- estimating
The work behind this page
Builds from our portfolio that this page draws on.
GroundUp
A construction project-management command centre for general contractors that keeps schedule, RFIs, budget and the field log in one place — and maps the critical-path recovery the moment a job slips.
Real EstateOpen Vision PPE Monitoring
Boundary surveillance, PPE compliance monitoring, and intrusion detection via real-time video analytics. Runs fully on-premise — no cloud required.
Safety & ComplianceRead next
- Extraction is clean on the issued PDFs and useless on the scanned onesMeasure effective resolution at the smallest annotation, the skew angle across the sheet, and the contrast between linework and background. Those three explain almost every scan that will not read.diagnostic
- Two estimators take off the same set and return different quantitiesNeither of them is careless. They measured different things, because nobody wrote down where one condition stops and the next begins — and comparing totals hides exactly that.diagnostic
- A takeoff condition is a measurement rule, not a highlighted areaThe coloured region on the sheet is the output of a condition, not the condition itself. The condition is the record that says what was measured, in what unit, with what adjustments, and where the number goes next.definition
- Bid alternates are several estimates wearing one cover sheetAn add or deduct alternate is a second estimate, not an adjustment. It has its own quantities, its own sub scope, its own duration and its own general conditions, and it has to survive into award as its own object.definition
- Every length in the takeoff is wrong by the same percentageCounts are right, lengths are proportionally wrong, areas are wrong by the square of the same factor. That signature points at scale — and the ratio you measure tells you whether one calibration fixes it or nothing does.diagnostic
- The allowance line: what it is, and what it quietly defers to laterAn allowance is not a placeholder number. It is a decision somebody has agreed to make later, with a date attached — and the estimate rarely records either the decision or the date.definition
Working on something in this space?
Tell us where you are in a sentence or two. We'll tell you honestly whether we're the right team, and what a sensible first slice of the work looks like.
Start the conversation