Ask three finance leaders what it costs them to process an invoice and you will get three numbers that cannot be compared. One counts only the accounts payable salaries. One includes the ERP licence, the scanning contract and a share of the shared service centre's floor space. One quotes a figure from a benchmarking deck they saw at a conference and has never calculated it at all. This is the reason invoice processing benchmarks were simultaneously the most-cited and least-useful metric in the back office by 2013. Everyone knew the headline range — that best-in-class organizations processed invoices for a small fraction of what typical ones spent, with the spread commonly quoted as an order of magnitude. Almost nobody could say whether their own number was inside or outside it, because they had not defined what they were counting.
What the Benchmark Actually Measures
A defensible cost-per-invoice calculation has a numerator and a denominator, and both are contested. The denominator is easier but not trivial. Count invoices processed, not invoices received, and decide explicitly whether credit notes, employee expense claims, utility bills paid by direct debit and intercompany charges are in scope. A team that excludes credit notes and intercompany looks meaningfully better than one that includes them, and the difference is definitional rather than operational. The numerator is where comparisons die. A complete cost includes AP salaries and benefits, supervision and management, the allocated cost of shared service overhead, technology — ERP modules, scanning, OCR, workflow, e-invoicing network fees — facilities, and the often substantial cost of time spent by people outside AP: requisitioners chasing purchase orders, approvers reviewing invoices, procurement resolving price disputes, and treasury handling payment queries. That last category is the one that is nearly always excluded and frequently the largest single component. An organization that reduces AP headcount by pushing coding and chasing out to the business has not reduced its cost per invoice. It has moved the cost somewhere that the benchmark does not look.
| Boundary | Make the choice explicit |
|---|---|
| Transaction count | Processed invoices; decide scope for credits, expenses and intercompany charges |
| AP effort | Salaries, benefits, supervision and management |
| Shared costs | Technology, service overhead and facilities |
| Business-side effort | Requisitioners, approvers, procurement and treasury time |
Qualitative summary of this article's source text, not a measured outcome or performance estimate.
Why the Spread Between Best and Worst Is So Wide
The gap between a high-performing AP function and a typical one is large, and it is not primarily about how fast the clerks work. It comes from four structural factors. Touchless rate is the dominant variable. An invoice that arrives electronically, matches a purchase order and a receipt automatically, and posts without human intervention costs very little. An invoice that requires a human to open it, identify the supplier, find the purchase order, chase the receipt, code the lines, route it for approval and follow up twice costs many times more. The cost-per-invoice figure is largely a weighted average of those two populations, so the real question is what proportion of invoices are touchless. Purchase order compliance upstream determines exception volume downstream. Invoices without a purchase order cannot be matched and must be coded and approved manually. In organizations where a large share of spend arrives without a purchase order, AP is doing procurement's work at the end of the process, and no amount of AP efficiency fixes it. Supplier master data quality drives rework. Duplicate supplier records, wrong bank details, inconsistent names and missing tax registration numbers produce exceptions on every transaction with that supplier, indefinitely. And the arrival channel matters enormously. Paper invoices, PDF attachments to a shared mailbox, portal submissions and structured electronic invoices have very different processing costs, and most organizations run all four simultaneously without having decided to.
The Trap in Benchmarking
Benchmarks are useful for direction and dangerous as targets, for reasons that were already visible by 2013. Comparability is usually assumed rather than established. Cost per invoice varies legitimately with industry, invoice complexity, number of legal entities, number of currencies, regulatory requirements and the share of spend that is project-based or contract-based. A construction group with progress billing against contracts and a retailer with repetitive trade purchases are not doing the same job. The metric ignores quality entirely. An AP function can improve cost per invoice by approving faster with less checking. The cost shows up later as duplicate payments, missed early-payment discounts, overpayments against contract rates and, occasionally, fraud. Cost per invoice should never be read without payment accuracy, duplicate payment rate and discount capture alongside it. It also ignores relationships. Cost reduction achieved by paying late, disputing aggressively and being unreachable transfers cost to suppliers, who eventually price it back in or deprioritise you. And self-reported benchmark data skews optimistic. Organizations that participate in benchmarking studies are disproportionately those who believe they will do well, and the cost definitions they submit are frequently narrower than the study intends.
What to Do With a Benchmark
The organizations that got value from this exercise used the benchmark to find the gap and then ignored it while closing it. Calculate your own number honestly first, including business-side effort. A fully loaded figure that is embarrassing is more useful than a flattering one that excludes half the cost. Decompose by channel and by exception path. The average is not actionable. Cost per touchless invoice, cost per PO-matched invoice with an exception, cost per non-PO invoice — these tell you where the money goes and which change is worth making. Find the concentration. In most AP functions a small number of suppliers generate a disproportionate share of exceptions, and a small number of exception types account for most of the manual effort. That list is the improvement plan. Attack the upstream causes rather than the processing step. Purchase order compliance, supplier onboarding quality, contract price accuracy and invoice submission channel all determine AP's workload before an invoice reaches AP. Optimising the processing of an exception that should not exist is the least valuable available work. Track the metrics that constrain the cost metric. Duplicate payment rate, payment accuracy, discount capture, days payable outstanding and supplier query volume. A cost improvement that degrades any of these is not an improvement.
Practical Guidance for AP Benchmarking
- Define the numerator before you calculate anything. Include technology, supervision, overhead allocation and business-side effort, or accept that your number is not comparable to anyone else's.
- Measure the touchless rate as the primary operational metric. Cost per invoice is largely a function of it, and it is far more actionable.
- Decompose the average by channel and exception type. The blended figure tells you almost nothing about what to change.
- Look upstream for the causes. Purchase order compliance and supplier master quality set AP's workload; AP efficiency only affects what is left.
- Never report cost per invoice alone. Pair it with duplicate payment rate, discount capture and payment accuracy so that cost reduction cannot hide quality degradation.
- Identify the exception-generating suppliers. A short list of suppliers usually accounts for a large share of the manual effort, and most of it is fixable by conversation.
- Adjust benchmark comparisons for entity count, currency count and regulatory scope. Multi-entity, multi-jurisdiction operations genuinely cost more per invoice, and pretending otherwise produces targets nobody can hit.
- Re-baseline after any structural change. A new e-invoicing obligation, an acquisition or a shared service transition invalidates the previous comparison.
The Regional Adjustment
Benchmark figures published from North American and European data need real adjustment before they mean anything to a Gulf-based finance function, and the adjustments run in both directions. Entity multiplicity raises the floor. A regional group with mainland and free-zone entities across the UAE, Saudi Arabia, Qatar and Egypt runs separate ledgers, separate statutory reporting and separate tax treatment per entity. Intercompany volume is high, invoices frequently need to be routed to the correct entity before anything else can happen, and the per-invoice overhead is structurally above a single-entity comparator. Benchmarking against a single-country operation and concluding the team is inefficient is a common and unfair error. Bilingual documentation adds handling. Invoices arriving in Arabic and English, supplier names transliterated inconsistently between systems, and correspondence that must be issued in both languages all add effort that the benchmark definition does not contemplate. The supplier matching problem in particular — where the same entity appears under several transliterations — is a genuine driver of duplicate records and exceptions. E-invoicing has changed the baseline sharply. ZATCA requirements in Saudi Arabia moved invoicing to a structured, cryptographically stamped and, for the relevant phase, cleared format. That imposed implementation cost and ongoing compliance obligation — and it also delivered exactly the structured, validated, machine-readable invoice data that touchless processing requires. Organizations that treated it purely as a compliance project paid the cost and captured none of the efficiency; those that redesigned the AP process around it saw touchless rates move in a way that a decade of internal improvement projects had not achieved. Payment practice affects the surrounding metrics. Longer customary payment cycles in parts of the region, cheque usage in some sectors and the mechanics of the Wage Protection System for payroll all mean that days payable outstanding and payment channel mix look different here. Comparing those figures to Western benchmarks without context produces misleading conclusions in both directions.
What the Metric Looks Like Now
The underlying economics have shifted twice since 2013, and the benchmark has aged badly in an instructive way. Robotic process automation removed a layer of keystrokes. Intelligent document processing then removed much of the extraction and coding work that OCR could only partially handle. More recently, models that can read an unstructured invoice, infer the correct coding from historical patterns, identify the likely purchase order match and flag the genuinely ambiguous cases have made a high touchless rate achievable without the structured data that previously made it a prerequisite. The consequence is that cost per invoice is becoming a less interesting number. When the marginal cost of processing a clean invoice approaches zero, the average is dominated entirely by exceptions and by the cost of the errors that automation lets through at speed. The useful metrics are moving to exception rate, first-time-right rate, cost per exception and payment accuracy. There is also a new failure mode that the old benchmark cannot see. An automated system that codes confidently and incorrectly produces a clean-looking process with a quietly deteriorating general ledger. The control question is no longer how many people touched the invoice but whether anyone is checking a sample of what the machine decided — and how quickly a systematic misclassification would be caught. Which returns to the point that was true in 2013 and is more true now: the number is only worth having if you know precisely what is inside it, and it is only worth improving if you improve the things that generate the work rather than the speed of the people absorbing it.
Common Questions
Why are invoice processing benchmarks so hard to compare?
Because there is no standard definition of the cost base. Some organizations count only AP salaries; others include technology, supervision, overhead and the substantial time spent by requisitioners and approvers outside AP. The last category is usually excluded and is frequently the largest component.
What actually drives the difference between high and low cost per invoice?
Touchless rate above all — the proportion of invoices that arrive electronically, match automatically and post without human intervention. Purchase order compliance, supplier master data quality and invoice arrival channel determine that rate before an invoice ever reaches AP.
What should be measured alongside cost per invoice?
Duplicate payment rate, payment accuracy, early payment discount capture, days payable outstanding and supplier query volume. Cost per invoice can always be improved by checking less, and those metrics are where the consequence appears.
Do Gulf-based finance functions benchmark differently?
Yes. Multiple legal entities across several jurisdictions, high intercompany volume, bilingual documentation and transliterated supplier names all raise the structural cost per invoice. E-invoicing mandates have simultaneously added compliance cost and created the structured data that makes high touchless rates achievable.
AP Benchmarking Assessment — Outpace calculates your real cost per invoice, decomposes it by exception path, and identifies the upstream changes that move it.
