The pilots that started in January are graduating this month, and the request arriving on finance desks has changed shape. It is no longer for a tool that drafts. It is for permission to let the thing act — post the journal, send the chase, release the match, issue the credit. That request is usually framed as a technical milestone. It is not. Moving an AI agent in the back office from assistance to autonomy is an authority decision, and it belongs to the same people who approve spending limits and bank mandates.
The question is not whether the agent can do it. It is whose authority it is acting under, and who answers when it acts
Capability arrived faster than anyone's governance. The gap between those two things is where the next twelve months of incidents will come from.
Three grants, routinely collapsed into one
Autonomy is not a dial. It is three separate permissions that organisations keep bundling. The right to read — to see records, documents and correspondence. The right to propose — to prepare an action that a person reviews. The right to commit — to make a change the business is bound by. Most governance arguments become tractable the moment these are separated. A great deal of value sits in read and propose, both of which are cheap to grant and easy to reverse. Nearly all of the risk sits in commit, and commit is the only one that needs a board-level conversation.
An agent needs its own identity
The most common shortcut in early deployments is to run the agent under a shared service account, or worse, under the credentials of the person who built it. Both destroy the two properties you will need first: attribution and revocation. Give every agent a named non-human identity with its own credentials, its own entitlements and its own log stream. When something goes wrong at eleven at night, the questions are which agent did this, under what permissions, and how quickly can it be switched off. None of those have answers if the agent has been wearing somebody's badge.
Write down the authority envelope
Five dimensions, one page, signed by a person with real authority. Which actions it may take. On which objects or accounts. Up to what value. Within what time window. And by what route each action can be reversed. Anything falling outside the envelope escalates rather than proceeds. That sounds obvious and it is the clause most often missing, which is why agents in production tend to behave well until they meet a case nobody anticipated.
Reversibility, not accuracy, should set the limit
The instinct is to grant autonomy in proportion to measured accuracy. That is the wrong axis, because accuracy is an average and the damage is a specific event. Sort actions three ways. Reversible — a draft, an internal reclassification, a queue assignment. Reversible at a cost — a posted journal, a sent email, an issued credit note. Irreversible — money leaving the bank, a statutory filing, an acceptance that forms a contract. Grant broad autonomy over the first category immediately. Grant the second with thresholds and sampling. Treat the third as human by default and justify every exception individually. An agent with ninety-nine per cent accuracy and the authority to pay suppliers is a worse design than one with ninety per cent accuracy that can only prepare the payment run.
| Action class | Examples in the article | Suggested boundary |
|---|---|---|
| Reversible | Draft, internal reclassification, queue assignment | Broader autonomy within an approved scope |
| Reversible at a cost | Posted journal, sent email, issued credit note | Thresholds and sampling review |
| Irreversible | Bank payment, statutory filing, contractual acceptance | Human by default; justify exceptions individually |
Qualitative summary of this article's source text, not a measured outcome or performance estimate.
Agents have managers
Every agent in production needs a named owner in the business function it serves — not in IT — who is answerable for what it produces, reviews its exceptions on a stated cadence, and has the authority to suspend it without a change request. If nobody's objectives include the agent's performance, its errors will accumulate in a queue that everyone assumes somebody else is reading.
Three gates to production
Rather than a schedule, use criteria. Gate one, shadow to propose: the agent's proposed action matched what the human did in an agreed proportion of cases across a full cycle, including a month-end. Gate two, propose to bounded autonomy: reviewers have stopped changing its output, measured as an override rate rather than a feeling, and the reversal path has been tested for real. Gate three, widening the envelope: the rate of genuinely novel cases has flattened. If the agent is still meeting situations it has never seen, the envelope is premature regardless of accuracy.
Practical Guidance for AI Agent Back Office Design
- Separate read, propose and commit as three distinct grants.
- Issue each agent its own identity, never a shared or borrowed one.
- Write a one-page authority envelope and have it signed.
- Set autonomy by reversibility, not by accuracy scores.
- Name a business owner with power to suspend without a change request.
- Test the reversal path before granting the action, not after.
- Define the stop control: who, how fast, and what happens in flight.
- Review the exception queue weekly with the owner present.
The Regional Angle
The first thing that catches groups here is that delegated authority is a formal, documented instrument rather than an internal convention. Authorised signatories are registered, bank mandates name specific individuals with specific limits, board resolutions define who may bind the company, and powers of attorney are notarised documents with scopes written into them. None of these instruments contemplates a non-human actor, and an agent operating inside a person's delegation is, in substance, sub-delegating an authority that was granted personally. The practical consequence is a routing question: the agent's authority envelope should be approved through the same governance that approves the delegation of authority matrix — a board or committee decision minuted alongside spending limits — rather than through IT change control. That single choice converts an awkward audit conversation into a documented one, and it also forces the useful discipline of expressing the envelope in the same language as your existing limits. The second is coverage, which is unusually fragmented across this region and directly affects how much autonomy is safe. A group with operations in the Gulf, North Africa and Europe is running at least three different working weeks and three sets of public holidays that do not align in any given month. An agent escalating at four in the afternoon on a Thursday may be escalating into a queue nobody opens until Sunday, and the in-flight item sits for three days with a payment term running against it. The design response is counter-intuitive but correct: reduce autonomy outside covered hours rather than increasing it. Define escalation coverage explicitly by jurisdiction calendar, have the agent hold rather than act when no reviewer is available for a class of decision, and monitor the age of the escalation queue as a first-class operational metric. Continuous operation is only an advantage when somebody is awake to receive the exception. The third is insurance, and almost nobody has asked the question. Commercial crime and fidelity cover is written around employee dishonesty and third-party computer fraud, with wordings placed through local brokers on international forms and endorsed for regional use. An erroneous payment initiated autonomously by software your own organisation deployed sits awkwardly between the insured categories — it is neither an employee's dishonesty nor an outsider's intrusion. Before granting any agent authority over outbound payments, put the scenario to your broker in writing and get the answer in writing. It costs an email, the answer may change your design, and finding out after an incident that the loss falls outside cover is a materially worse way to learn it.
The objection worth taking seriously
The strongest objection is proportionality. Most of what these agents do is draft correspondence, sort queues and prepare reconciliations. Building an identity model, an authority envelope, a gating process and a governance route around a tool that mostly writes emails is exactly the kind of overhead that makes internal IT slower than the market, and while one organisation is drafting its envelope, its competitor has switched the thing on and is three months ahead. That is fair for the read and propose grants, and the response is to grant those quickly and with minimal ceremony. Nobody needs a board paper to let software draft a supplier chase. The case for the discipline is entirely about the commit grant, where the cost is asymmetric. A well-run agent produces a steady stream of small savings; a single wrong irreversible action produces an incident, an audit finding, and in most organisations the immediate suspension of every agent in production, including the ones that were working. The envelope is one page and takes an afternoon. The identity model is an hour of configuration. Both are written once and reused for every subsequent agent, which means the second deployment carries almost no governance cost at all. Framed that way it is not bureaucracy; it is the thing that lets you keep saying yes as the requests multiply through the year.
Common Questions
Can one agent hold different authority in different entities?
It should. Entitlements belong to the identity, and a group with varied delegation matrices needs the envelope expressed per entity rather than globally.
What override rate indicates readiness?
Less important than the trend and the reason. A stable low rate with reviewers who can no longer explain why they changed something is the real signal.
Should the agent be allowed to email external parties?
Treat outbound external communication as reversible at a cost, not as reversible. Sampling and tone constraints, plus a hard exclusion list of counterparties in dispute.
What should we expect over the next twelve months?
Expect autonomy requests to arrive faster than governance can absorb them, and expect the pressure to come from operations rather than technology. Expect auditors to begin asking who authorised a non-human actor and under what instrument, which is a question most organisations cannot currently answer. Expect identity platforms to ship non-human identity management as a headline capability during the year, because the gap is now obvious. And expect at least one public incident involving an autonomously initiated payment — after which every envelope in the market will be rewritten, and the organisations that already have one will simply tighten a threshold.
AI Agent Back Office Design — we separate read, propose and commit, write the authority envelope your auditors will ask for, and route it through the governance that already approves your limits.
