Every automation vendor has spent the first quarter of this year relabelling its product. Screen-scraping robots built in 2019 are now described as agents, orchestration engines have become agentic platforms, and the renewal conversation arrives with a new slide deck and the same licence count. Some of it is genuine. Agentic automation is a real architectural shift away from rules-based bots, and it changes what can be automated in a back office. It also changes how automation fails, which is the part the slide deck omits.
A rules-based bot fails the same way every time. An agent fails a different way every time — and that single difference is what your operations model has to absorb
Everything else in this piece follows from that sentence.
What robotic process automation was actually good at
The criticism of the last decade's bots has become lazy. They were brittle, yes. But they were also deterministic, cheap to certify, and easy to explain to an auditor: here is the sequence, here is the log, it does this every time. Most importantly, when a bot broke, it broke identically and visibly. Somebody fixed the selector and it ran again. That is a genuinely valuable property and organisations are about to discover how much they liked it. What rules could not do was handle variance. Every exception needed a new branch, the branches multiplied, and eventually the maintenance cost of the decision tree exceeded the cost of the people it replaced.
What agentic actually means
Strip away the marketing and an agent is a loop. It observes a situation, plans a next step, calls a tool, reads the result, decides whether it is finished, and repeats. The new capability is not that it can write fluent text. It is that the system chooses the sequence of actions rather than following one you specified. That is the whole difference, and it is both the reason agentic automation can handle the long tail and the reason it needs an engineering discipline that bots never required.
Five properties you now have to engineer
Idempotency on every action. An agent may retry a step it believes failed. A duplicated payment is not a retry. Every action with a side effect needs a key that makes repetition harmless. A bounded action space. Not "access to the browser" but an explicit, short list of tools with typed arguments. The safety of the system is the size of that list. Step and cost budgets. Loops are unbounded by default. Cap iterations, cap spend, and make exceeding the cap an escalation rather than a crash. Side effects behind a transaction boundary. The agent should propose; deterministic code should commit. Nothing irreversible happens inside the planning loop. A replayable trace. "Why did it do that" is now a question somebody will genuinely ask, and the answer has to be reconstructable from stored reasoning, tool calls and inputs. None of these were necessary with rules-based automation. All five are mandatory now, and the pilots that skip them are the ones that produce an incident rather than a case study.
| Boundary | Control |
|---|---|
| Repeated action | Use idempotency keys for side effects |
| Available actions | Expose a bounded, typed tool list |
| Planning loop | Cap steps and spend; escalate on breach |
| Irreversible commit | Let deterministic code enforce the transaction boundary |
| Later review | Retain inputs, tool calls, results and a replayable trace |
Qualitative summary of this article's source text, not a measured outcome or performance estimate.
Where each approach actually wins
Rules win where volume is high, the path is stable and the process is regulated. Payroll runs, statutory filings, standard payment batches. Do not rewrite these. They work, they are certified, and variance is the enemy rather than the opportunity. Agents win in the long tail: supplier queries that arrive as free text, exception triage where the next step depends on what you find, reconciliation of mismatches, chasing missing documentation, and the enormous category of work that was never automated because writing the rules would have cost more than doing it manually. The architecture that survives production is a hybrid. The agent reads, decides and routes. Deterministic code executes. The agent is the judgement and the explanation; the existing automation is still the hands.
Testing has to change shape
You cannot write a conventional test suite for a non-deterministic planner, and pretending otherwise is how pilots pass and deployments fail. Test at the boundary instead: given this input, did the agent call the correct tool with the correct arguments? That is deterministic and assertable. Then evaluate in aggregate against a scored scenario set rather than expecting identical outputs. Then run in shadow mode against live traffic for several weeks, comparing proposed actions against what humans actually did, before anything is allowed to commit.
Practical Guidance for Agentic Automation Pilot
- Make every side-effecting action idempotent before the first pilot.
- Define the tool list explicitly; treat its length as a risk measure.
- Cap steps and spend, and escalate on breach rather than failing silently.
- Let the agent propose and deterministic code commit.
- Store the full trace — inputs, reasoning, tool calls, outputs.
- Run shadow mode against live volume for at least a month.
- Target the unautomated long tail, not processes rules already handle.
- Keep your existing automation and call it from the agent.
The Regional Angle
The first constraint here is that a great deal of back-office work is conducted through government and bank portals that were never designed to be automated, and in several cases explicitly forbid it. Customs declarations, labour and immigration submissions, tax filings, wage protection uploads and bank transaction portals typically require a named individual's credentials, a one-time password sent to a registered mobile number, and occasionally a physical token. An agent operating those credentials is not just a technical workaround; it puts a named person's authorisation behind actions they did not review, which is a governance problem before it is a security one. Read the terms of use of each portal before scoping anything, separate read-only retrieval from submission, and where submission must remain human, design the agent to prepare and stage the work rather than to file it. That split usually captures most of the benefit and all of the defensibility. The second is physical. Post-dated cheques, stamped and signed hard-copy documents, original delivery notes and cash collection remain ordinary features of commerce across much of the region, and a substantial share of back-office exception volume is gated by an object that has to move between two hands. No amount of planning-loop sophistication touches it. Before sizing an agentic business case, measure what proportion of your current exception queue is blocked on a physical instrument or a wet signature rather than on a decision. In many regional finance operations the figure is high enough to halve the addressable benefit, and it is far better to know that during scoping than to explain it at the first quarterly review. The third is commercial, and it will quietly decide whether any of this scales. Most regional automation capability was built by implementation partners priced per process automated, which was rational for bots and is actively counterproductive for agents. The entire point of an agentic approach is that one well-designed capability generalises across many similar cases — which under a per-bot commercial model reduces the partner's revenue. Expect proposals that decompose an agent into fourteen chargeable automations. Renegotiate to an outcome basis tied to straight-through volume or cost per exception handled, or build the capability internally and use the partner for integration work. Organisations that leave the old pricing model in place will get the old architecture back, wearing new language.
The objection worth taking seriously
The strongest objection is memory. Organisations spent five years and considerable money building bot estates that now require a maintenance team, break whenever a vendor changes a screen, and delivered a fraction of the promised saving. The people being asked to fund agentic automation are frequently the same people who funded that, and they are entitled to ask what is different beyond the vocabulary. Adding non-determinism to a category of technology that already disappointed is not obviously an improvement. That scepticism is well earned and should be answered directly rather than deflected. The honest answer is that the failure mode is better matched to the work. Rules-based automation failed because real processes vary and rules cannot; every exception became a branch until the tree collapsed under its own weight. Agents handle variance natively and, in exchange, cannot guarantee identical behaviour. So the correct use is not to re-automate what the rules already do well — that trades a working certainty for an unnecessary probability. It is to reach the work nobody ever automated because the rules would have cost more than the labour. That is a genuinely new addressable area rather than a re-run. And the discipline is unglamorous: keep the bots, bound the action space, make everything reversible, and start where a mistake is cheap. Anyone selling this as a wholesale replacement for your existing estate is selling the previous cycle again.
Common Questions
Should we retire our existing bots?
No. Keep the deterministic paths and expose them to the agent as tools. The bots become the execution layer rather than the decision layer.
How autonomous should a first pilot be?
Proposing only. Let it draft the action and have a person commit it, for long enough to build a comparison set. Autonomy is earned with evidence.
What is the most common cause of pilot failure?
An unbounded action space. Agents given broad system access do something unexpected within weeks, and the resulting incident ends the programme.
What should we expect over the next twelve months?
Expect every established automation vendor to ship agent features into existing platforms during the year, which will make procurement easier and evaluation harder. Expect the first well-publicised incident involving an autonomous agent taking a financial action, and expect it to tighten governance across the market faster than any framework. Expect tooling for tracing and evaluating agent behaviour to improve quickly, because it is currently the weakest part of the stack. And expect the organisations doing well by year end to be the ones that spent this spring making their actions idempotent rather than choosing a platform.
Agentic Automation Pilot — we start where a mistake is cheap, bound what the agent can touch, and prove it in shadow mode before it commits anything.
