Operated function · document intake

An accuracy figure is useless unless it can tell you which ones it got wrong.

Right most of the time, with no way to know which times, means somebody still checks every field — so the saving is nothing. The question worth asking a vendor is not how accurate they are; it is what happens on the documents they are unsure about.

What arrives, and what somebody does with it

The inbound pile is never one format. It is a clean structured file from one supplier, a well-made PDF from another, a photograph taken on a phone at an angle, a scan of a document that was itself a fax, and a form somebody completed in handwriting and posted.

A person opens each one, finds the fields that matter, and types them into a system. The work requires almost no expertise and total attention, which is the worst combination available — errors come from the attention lapsing rather than from anybody not knowing what to do.

The errors are not evenly distributed. They cluster on the documents that were hard to read, on the fields that look similar to another field, and at the end of a shift. Nobody records which, so the pattern is invisible and the response is a general instruction to be careful.

Checking is done by sampling, because checking everything would cost as much as the original entry. So the accuracy figure for the whole population is inferred from a fraction, and the errors that were not sampled are found later by whoever depends on the data.

And exceptions consume the experienced people. Anything that does not fit the expected shape goes to whoever knows the edge cases, which means the most capable person spends the day on the least tractable ten per cent.

Accuracy without calibration means everything still has to be checked

An extraction system that cannot say which fields it is unsure about forces a full manual check regardless of its accuracy rate — so the headline accuracy figure has almost no bearing on the work saved.

Calibration is the property that matters and it is rarely sold, because it sounds like a weakness. A system that says "this field, I am confident; that field, I am not" lets a person check only the second, and that is where the actual saving comes from — not from being right more often, but from being right about when it is right.

So the design point is that uncertainty is declared rather than hidden. A low-confidence field is routed to a person, always, with the document image beside it. Nothing is written into a system of record from a guess, because a wrong value that entered silently is more expensive than a field that waited an hour.

The second point is that the error pattern must be recorded. Which fields, which document types, which sources — because that turns a general accuracy figure into a specific fix: a supplier whose format is registered once, a form field that is ambiguous and should be redesigned, a source that should be received structured rather than as a scan.

And the highest-return finding is usually upstream. A document that arrives as a photograph of a printed copy of a structured file could have arrived as the structured file, and asking the sender is cheaper than any extraction technology.

What is not delegated is the meaning. Whether an invoice should be paid, whether a form is complete enough to act on, and whether a value is plausible in context are judgements, and extraction is not judgement.

What moves, and how you would know

Fields a person actually has to check — measured by fields routed for review as a share of all fields, which is the honest measure of work saved rather than an accuracy percentage.

Whether confidence predicts correctness — measured by error rate within high-confidence fields versus low-confidence ones — calibration, which is the property that decides everything else.

Errors reaching a system of record — measured by incorrect values written, found by downstream reconciliation, against your own baseline.

Where errors cluster — measured by error rate by field, document type and source, which converts a general figure into a specific upstream fix.

Documents that could have arrived structured — measured by the count by sender, which is usually the cheapest available improvement and requires no technology at all.

Experienced staff time on exceptions — measured by hours on non-standard documents, sampled the same way before and after registration of the recurring formats.

a headline accuracy percentage, deliberately. It is the number the category is sold on and it is close to meaningless without calibration — an uncalibrated system at any accuracy still requires a full check. Nothing here decides whether to pay an invoice, accept a form, or act on what a document says.

Into the system of record, from wherever documents arrive

Extracted values write into your existing system through its documented interfaces, with the source document and the per-field confidence recorded alongside. Nothing creates a parallel data store — the point is to populate the system you already work from.

The format registry is yours: which sender, which document type, which fields, where they sit. It is written, versioned and owned, and it is the asset that makes recurring documents cheap. It also survives the engagement.

Documents stay in your document store. A processing operation that accumulated its own copy of every invoice and form you receive would have created a second disclosure surface for no benefit.

Writing into a system of record

Nothing is written from a guess. A field below your confidence threshold goes to a person with the document image beside it, and the threshold is set by you rather than tuned to maximise automation rate — a threshold tuned for coverage is a threshold that writes wrong values.

Every written value carries its source: which document, which page, which region, and the confidence at the time. When an error is found later, every value from the same source and pattern is findable rather than estimated.

Documents frequently contain more than the fields being extracted — bank details, personal information, commercial terms. Access is scoped to the operation and no document leaves your store.

Operational access is not permission to train. Your documents and the data extracted from them do not become material improving anything serving another organisation, including your suppliers or competitors whose terms appear in them.

Finance or operations control, and whoever owns the system of record

The control question is what may be written automatically and what must be reviewed, and it is a threshold your own control function sets. A vendor-set threshold optimises for the vendor’s automation-rate metric, which is not your control objective.

Whoever owns the receiving system should confirm the write path and the audit fields, because a value arriving with no recorded source is a value nobody can investigate later.

Where an obligation attaches through a data class or a retention schedule, it is marked applicability-gated rather than presented as standing.

Test calibration on documents whose answers you already know

A sample of already-processed documents where the correct values are known — read-only, nothing written — measuring not accuracy but whether confidence predicts correctness.

The test that matters runs against documents you have already keyed, so the right answer exists. It produces two numbers: how often high-confidence fields are correct, and how often low-confidence fields are wrong. If those separate cleanly, the system is calibrated and the review burden can shrink. If they do not, the accuracy figure is irrelevant and you should not proceed.

The same pass produces the error-cluster map and the count of documents that could have arrived structured. The second is frequently the largest single saving and requires asking a sender rather than buying anything.

If you continue, the first delegation is one registered document type from one sender, with your threshold set, uncertain fields routed to your people, and nothing written that was not above the line.

Questions buyers actually ask

We were quoted higher accuracy by another vendor.

Ask them which fields they got wrong, and what happens on the ones they are unsure about. An uncalibrated system at any accuracy still requires somebody to check everything, because there is no way to know which items to check — so the headline number does not translate into work saved. The calibration test runs on documents you have already keyed, costs a read-only pass, and settles the comparison on your own documents rather than on a benchmark.

Our documents are messy — photographs, scans of faxes, handwriting.

Then a meaningful share will fall below any sensible threshold and route to a person, which is the correct outcome rather than a failure. The value in a messy population is often less in extraction and more in the two things the first pass produces: which senders could be sending structured data instead, and which fields are ambiguous enough to be redesigned. Both are upstream fixes worth more than any extraction rate.

A wrong value in our system of record is a serious problem.

Which is exactly why nothing is written below your threshold and why the threshold is yours rather than ours. It is also why every written value carries its source document, region and confidence — so when an error is found, every value from the same pattern is findable rather than estimated. If your control position is that no value may be written without human confirmation, the operation still works: everything routes for review, and the saving is in the finding and presenting rather than the writing.

This is data entry. We can offshore it more cheaply.

Often true on unit cost, and it does not address the two structural problems: errors that are invisible until somebody downstream finds them, and the same recurring formats being re-keyed forever because nothing registers them. If cost per document is your only constraint, the comparison is legitimate and may well favour offshoring. The argument here is about calibration and about upstream fixes, neither of which a cheaper keyboard provides.