Five days after a markedly more capable language model was released, and three days after Microsoft showed an assistant sitting inside the productivity tools most finance teams live in, every accounts payable vendor's roadmap has acquired an agent. Some of those agents are running in production somewhere today. Most were demonstrated on a stage last quarter and will ship in a release named after a season. It is worth separating the two, because accounts payable is one of the few back office functions where the difference between a suggestion and an action is measured in money leaving the building.
Reading the invoice was never the hard part. Deciding what to do about the one that does not match is the hard part
Document capture has been competent for years. In a well-run payables operation the majority of invoices already arrive structured or are extracted accurately, match a purchase order and a receipt, and post without a person seeing them. That portion of the process is not where the cost sits. The cost sits in the remainder — the invoice with no purchase order, the quantity that differs by two units, the supplier who invoices in a different currency from the order, the service charge nobody can allocate. Each of those requires someone to investigate, decide and often ask a question of a person outside finance. That is where the hours go, and it is the only part of payables where an agent is interesting.
Four rungs, and where the market actually is
Extract. Mature, unglamorous, largely solved for common document types. Classify and code. Suggesting the account, cost centre and tax treatment from history. Probabilistic, genuinely useful where volumes are high and coding patterns are stable. Decide and act within bounds. Holding an invoice, routing it to a named approver, requesting a goods receipt, flagging a probable duplicate. This is the new capability and the reason the word agent appears. Converse. Handling the supplier's email about the unpaid invoice, or the budget holder's question about why something was rejected. Most of what is being sold as agentic in payables this quarter is the second rung with a chat interface attached. That is not a criticism — the second rung has real value — but it should be priced as what it is.
What the early deployments actually show
The reported gains cluster in exception handling rather than in straight-through rates, which is the opposite of what the marketing implies. Teams running these systems describe shorter investigation time per exception, faster responses to supplier queries, and fewer invoices sitting untouched because nobody knew who owned them. The straight-through percentage barely moves, because the invoices that already matched were never the problem. The second consistent finding is that the last step stays human. The agent assembles the case, drafts the query, proposes the coding — and a person presses send or approves the post. Deployments that removed that step early are the ones being quietly rolled back.
Three failure modes already visible
Plausible miscoding. An unusual invoice gets coded confidently to the cost centre that resembles it most. Nobody queries a suggestion that looks reasonable, and the error is found at year end by someone reconciling a budget. Duplicate exposure. An agent resolving a mismatch by locating a second matching document can create the conditions for a duplicate payment rather than prevent one, particularly where the same invoice arrived twice through different channels. Unintended commitment. A supplier-facing message that states when payment will be made. It is a short sentence, it is helpful, and in some jurisdictions it is evidence.
Design the boundary around reversibility, not around confidence
The useful control question is not how accurate the model is. It is what happens if this particular action is wrong. Let the agent act freely on reversible things: asking a question, requesting a document, routing for review, drafting a response, proposing a code. Require human approval for anything irreversible: posting to the ledger, releasing a payment, closing a query, amending a purchase order. And put one item entirely outside the boundary. No automation, agentic or otherwise, should be able to create or amend supplier bank details. That single field is the target of most payment fraud, the change request usually arrives looking entirely legitimate, and the control that protects it is a call-back to a number held on file. An agent processing that change efficiently is an agent doing the fraudster's work at machine speed.
Piloting this without a mess
Start with exceptions rather than the happy path, because the happy path is already automated and improving it proves nothing. Establish a baseline first: average touch time per exception, number of exceptions aged over ten days, supplier query response time. Without those three numbers measured before you start, the pilot will be evaluated on impressions. Run the agent inside the payables system's own permission model rather than beside it, so that everything it does is attributable, logged and visible to the same people who review human activity. Require that every action it takes is reviewable as a normal audit entry, not as a log file somebody has to request. And agree the stop rule in advance — the specific outcome that ends the pilot — because pilots without one become permanent by default.
Practical Guidance for AP AI Pilot Design
- Aim at exceptions, not at invoices that already match.
- Baseline three numbers before the pilot starts.
- Bound actions by reversibility, not by model confidence.
- Exclude supplier bank detail changes from any automation.
- Keep the agent inside the system's permission model.
- Log every action as an auditable entry, reviewable by finance.
- Match suppliers on registration number, not on name similarity.
- Agree the stop rule before the first invoice is processed.
The Regional Angle
Three factors change the calculation for payables teams in this region. The first is that the clearance model has already solved part of the problem that vendors are now selling AI to solve. Where electronic invoicing operates on a clearance basis, as the Saudi regime does with its phased integration waves, an invoice arrives validated by the tax authority and cryptographically stamped. Its structure, its arithmetic and the identity of the issuer have been checked before it reaches you. Spending an automation budget on extracting and validating fields from that document is spending money on a solved problem. The risk has moved, and it has moved somewhere less convenient: whether the goods or services were actually received, whether the purchase was authorised in the first place, and whether the supplier is one you intended to trade with. Regional teams should therefore point their pilot at receipt matching and supplier verification, and treat document reading as infrastructure they already have. The second is a matching problem that imported products handle badly. Invoices here are routinely bilingual, supplier legal names are transliterated inconsistently across systems, and the same entity appears as three records with three spellings. Any agent that matches or de-duplicates on name similarity will do one of two damaging things: merge two genuinely different suppliers because their transliterations converge, or fail to spot a duplicate invoice because the same company was entered twice with different spellings. The fix is a data decision rather than a model decision. Make the tax registration number the primary matching key, require it on onboarding, backfill it for active suppliers before the pilot, and treat name matching as a hint that triggers review rather than as a conclusion. The third is about how much of the process is actually in the system. In a great many regional finance functions, the invoice sits in the payables platform while the decision about it happens in a messaging app, a phone call or a corridor. The approval recorded in the system is the ratification of a decision made elsewhere, and cheque runs and manually instructed transfers remain common. An agent placed on top of that process automates the documented half of something whose real decisions occur off-system, and it will produce a beautiful audit trail of a fiction. Before piloting, measure the share of payment decisions with a genuine in-system approval, made by the person who actually decided, before the payment was instructed. If that share is low, fix it first. It is less exciting than an agent, and it is worth more.
The objection worth taking seriously
The strongest objection is that payables is precisely the wrong place for probabilistic technology. It is a rules-governed, audited, transaction-safe function where deterministic automation already works: robotic process automation and configured matching tolerances handle high volumes predictably, produce identical results for identical inputs, and can be explained to an auditor in a sentence. Introducing a model whose output varies, cannot be fully reproduced and cannot explain itself adds variance to a controlled process in exchange for savings in a function that is not usually the largest cost in the business. That argument is strong, and the auditability concern in particular is not adequately answered by anything on the market this month. Where it falls short is in what the deterministic tools actually cover. Rules handle the invoices that follow the rules, which is exactly the population that was never expensive. They break at the exception, and at that moment the process hands over to a person who reads three systems and sends two emails. The case for a model is not that it should replace deterministic matching — it should not — but that it can compress the investigation a human was already performing using judgement, with that human still approving the outcome. Keep the determinism where determinism works, keep the approval gate exactly where it is, and apply the model only to the part that was already discretionary.
Common Questions
Should this replace our existing capture tool?
Almost certainly not. Capture is a solved, cheap layer. The agent sits above it, and replacing a working extraction pipeline to buy an assistant is a costly way to get a chat box.
How do we explain model-assisted coding to auditors?
By keeping a human approval on every posting, logging the suggestion and the approver separately, and being able to show the population where the suggestion was overridden. If you cannot produce that override rate, you are not ready.
What volumes justify this?
Look at exception count rather than invoice count. A team handling two thousand invoices a month with four hundred exceptions has a stronger case than one handling ten thousand with fifty.
What should we expect over the next twelve months?
Expect every finance and payables vendor to announce an assistant this year, accelerated by last week's releases, and expect most of those announcements to describe coding suggestions rather than autonomous action. Expect external auditors to begin asking structured questions about model-assisted coding within about a year, which argues for keeping override statistics from the start. Expect confidentiality and data location terms, rather than accuracy, to be the blocker for regional buyers, particularly where invoice images contain commercially sensitive pricing. And expect genuine autonomy to stay confined to reversible actions well into next year, because the first publicised case of an agent releasing a fraudulent payment will set the market back further than any capability gain moves it forward.
AP AI Pilot Design — we baseline your exception handling, draw the boundary where reversibility actually changes, and make sure the pilot proves something your auditor can live with.
