Intelligent document processing has been promised to back-office teams for about fifteen years, and for most of that period the promise was false. What changed is not marketing. It is that the underlying approach was replaced. The previous generation worked by templates. You told the software where on the page the invoice number sat, where the total sat, where the tax figure sat, and it read those coordinates. One template per supplier layout. A supplier redesigned its invoice and the template broke silently. The current generation reads a document the way a person does — jointly interpreting text, position and visual structure — which means it can find a total it has never seen in that position before. Cloud services now ship prebuilt extractors for invoices, receipts and identity documents that require no training at all, and custom models that need a handful of examples rather than a library of layouts. That is a genuine architectural shift, and it is why the technology has moved from perpetually-piloted to routinely deployed in the last eighteen months. It is also why most organisations are about to buy it badly.
The machine's job is not to read the invoice
Here is the reframing that determines whether a deployment succeeds. The objective is not extraction accuracy. The objective is triage: the software's real job is to decide which documents a person must look at, and to be right about that decision. A pipeline that reads ninety-five per cent of fields correctly but cannot tell you which five per cent are wrong has produced nothing of value, because a human must still check everything. A pipeline that reads ninety per cent correctly and reliably flags the documents where it is uncertain has removed ninety per cent of the work. Accuracy matters less than calibrated self-doubt.
Three numbers, and the one that gets hidden
Straight-through rate. The proportion of documents that pass from arrival to posted with no human touch. This is the only number that maps to money. Field-level accuracy on the fields you post. Not character accuracy, not average accuracy across thirty extracted fields. Accuracy on supplier identity, invoice number, date, net, tax and total — measured per field, because a ninety-nine per cent character rate can still mean one invoice in five has a wrong digit somewhere that matters. Cost per document, fully loaded. Licence plus per-page fees plus review labour plus the amortised setup. And the number vendors do not put on slides: exception-handling time. Nearly every deployment makes the exceptions harder, because the easy documents go straight through and the residue is uniformly awkward. If your reviewers previously processed a mixed stream at two minutes each and now handle only the difficult ones at six minutes each, your saving is real but smaller than the demonstration implied. Measure it before and after.
Confirm entity details
Check supplier and buyer identifiers against controlled records.
Match supporting evidence
Compare the order and goods receipt where applicable.
Check arithmetic and duplicates
Validate totals and potential duplicate combinations.
Route exceptions
Keep uncertain or mismatched records for review.
Qualitative summary of this article's source text, not a measured outcome or performance estimate.
Validation beats extraction
The largest accuracy gains in any live deployment come not from the model but from business rules applied to what it produced. A guess that passes validation is worth more than a read that was never checked. Match against the purchase order and the goods receipt. Identify the supplier by tax registration number rather than by name. Recompute the arithmetic — line extensions, subtotal, tax at the applicable rate, total — and reject anything that does not foot. Check for duplicates on the combination of supplier, invoice number and amount, and again on amount and date, because the most common payables error in the world is paying the same invoice from an emailed copy and a posted original. Check that the tax treatment is consistent with the supplier's registration status. Each of those rules converts an uncertain extraction into either a confident posting or a flagged exception. That is the mechanism by which straight-through rates move from sixty per cent to eighty-five, and none of it is machine learning.
The review screen is the product
Teams evaluate extraction engines and then hand reviewers an interface that shows the document on one side and a form on the other, with no indication of where each value came from. The reviewer, sensibly, re-reads the whole document. The automation has saved nothing and added a screen. The review experience needs three properties: every extracted value highlights its source location in the image when selected; correction is keyboard-first, field to field, without a mouse; and corrections feed back into the model or the rules rather than dying in a session. Ask to see the review screen before you ask about accuracy, and ask how many seconds a trained reviewer needs per exception. If the vendor has not measured that, they have not deployed at scale.
Where to start, and what to leave alone
Start with the highest-volume, most structured, least consequential document type you have, which for almost everyone is supplier invoices matched to purchase orders. Then delivery notes and goods receipts. Then expense receipts. Then inbound customer purchase orders. Leave contracts, handwritten forms and anything requiring judgement for later. They are the documents people demonstrate because they are impressive, and the documents that fail in production because there is no validation rule to catch a wrong answer.
Practical Guidance for Document Processing Pilot
- Baseline first — current cost per document, touch time, error rate and rework, measured before you see a demonstration.
- Run the pilot on five hundred of your own documents, in their real distribution, including the worst suppliers.
- Define straight-through criteria explicitly, field by field, with confidence thresholds set higher for amounts than for descriptions.
- Build the validation rules first; they produce more benefit than a better model.
- Identify suppliers by tax registration number, never by name matching.
- Time the exception path before and after, and staff it deliberately.
- Insist on a source-highlighted, keyboard-first review screen and measure seconds per exception.
- Keep contracts short and per-document pricing transparent, because unit costs in this market are falling.
The Regional Angle
Three regional document problems are worth more than the invoice pilot everyone starts with. The first is proof of delivery. Across regional distribution, the evidence that goods arrived is a paper delivery note signed and stamped by a storekeeper at a customer's gate, carried back by a driver, and filed — sometimes weeks later. When a customer disputes an invoice or withholds payment, that stamped note is the entire receivable. The input to any capture process is therefore a photograph taken on a driver's phone in bright sun, at an angle, possibly with a thumb across the corner. This is a solvable problem, but only if capture happens at the gate with an immediate automated quality check that tells the driver to retake the picture, rather than through batch scanning at the end of the month when the truck has done forty more stops. Organisations that fix this recover money they were previously writing off, which is a more compelling first case than shaving seconds off invoice entry. The second is the buyer block, and it silently destroys tax recovery. Suppliers in this region address invoices to whichever version of the group name they happen to know, and a group with eleven registered entities across onshore and free-zone jurisdictions will receive invoices naming the wrong one constantly. Tax recovery depends on the document naming the correct registered entity with the correct registration number. Almost every extraction pipeline validates the supplier's details thoroughly and accepts the buyer's details without checking, which means an automated payables process will cheerfully post and pay irrecoverable input tax at scale and at speed. Validate the buyer block against your own entity register, and treat a mismatch as an exception requiring a corrected invoice rather than a note in the comments field. The third is the document regional finance teams actually spend their week on, which is not the invoice. It is the supplier statement of account, arriving monthly as a PDF, and the customer ledger reconciliation arriving by email from the other side. Extracting statements and reconciling them automatically against your ledger finds missing invoices, unapplied credit notes, wrongly allocated payments and disputed balances — the unglamorous arithmetic that consumes the first ten days of every month and occasionally reveals that a supplier has been charging for something delivered twice. The document is highly structured, the volume is high, and almost nobody automates it because the invoice is the more obvious target.
The objection worth taking seriously
The strongest objection is that this whole category is a sunset technology. The best document automation is not receiving a document at all, and mandates are steadily making that happen. Italy has required structured electronic invoicing for years. Saudi Arabia's generation phase began in December and system-to-system integration starts in January. Other jurisdictions are drafting. When an invoice arrives as validated structured data, extraction is not merely unnecessary — it is a category error. On this view, building a pipeline to read PDFs in 2022 is automating a format regulators have already decided to abolish, and a three-year contract will outlive its own rationale. That argument is strong and largely correct about direction. Mandates achieve in five years what no software vendor can achieve at all. But coverage is partial and slow. Cross-border suppliers, non-resident vendors, free-zone entities and small trade suppliers sit outside the covered population for years, and payables are not organised by jurisdiction. More importantly, invoices are a minority of the documents a back office handles: delivery notes, statements, customs paperwork, claims, certificates, remittance advices and contracts have no mandate behind them and will stay unstructured indefinitely. And the capability that actually produces the savings — validation rules, confidence thresholds, exception triage, a fast review path, feedback loops — transfers directly to the structured channel, because a perfectly structured invoice still has to be matched, validated, approved and posted. So the sensible posture is to buy the extraction layer thin, on short terms and transparent per-document pricing, and to point the pilot at the document types no mandate will ever reach.
Common Questions
Prebuilt or custom models?
Prebuilt for invoices, receipts and identity documents — they are trained on far more data than you will ever have. Custom models only for genuinely proprietary forms, and then with a handful of examples rather than hundreds.
What straight-through rate should we expect?
Sixty to seventy per cent in the first months for purchase-order-matched invoices, rising into the eighties once validation rules mature. Anyone quoting ninety-plus at the outset is describing a clean document set.
Do we still need people?
Yes, fewer and differently deployed. The work shifts from keying to exception judgement and supplier remediation, which is more skilled and harder to staff cheaply.
What should we expect over the next twelve months?
Expect prebuilt extraction quality to keep improving and per-document prices to keep falling, which argues for short contracts. Expect vendors to reposition from optical character recognition to document understanding, with the same products behind the new label. Expect the electronic invoicing mandate map to expand, with the Saudi integration phase from January the nearest deadline for regional groups. Expect the most valuable regional deployments to be in documents nobody markets — delivery notes, statements, customs packs. And expect published accuracy claims to remain unaudited, which means the pilot on your own documents is not diligence theatre; it is the only evidence there is.
Document Processing Pilot — we baseline your real cost per document, build the validation rules that make extraction safe, and prove a straight-through rate on your own worst suppliers before you sign anything.
