Internal control frameworks were designed around a specific assumption: that a person performs a step, another person checks it, and the evidence that both happened is a signature, an approval record or an initialled reconciliation. Segregation of duties, authorisation limits, review and approval — all of it rests on there being two people whose interests differ. When a process is operated by software that reasons rather than follows rules, that architecture does not fail dramatically. It fails quietly, by continuing to produce evidence that no longer demonstrates anything.
Segregation of duties assumes two parties with different incentives. Two agents from the same vendor running the same model are not two parties, and an approval one of them gives the other is not a control
This is the year finance and audit functions have to redesign around that, because the deployments are arriving faster than the frameworks.
The four controls that stop working
Segregation of duties. Splitting preparation from review between two automated steps provides no independence whatsoever. If the same model reasoning from the same data prepares and reviews, a systematic error passes both times. Independence now has to come from a different source — a deterministic rule, a different vendor's system, or a human. Review and approval. A human approving two hundred machine-prepared items per hour is not reviewing. The control is nominally present and substantively absent, and it is worse than no control because it creates documentary evidence of diligence that did not occur. Authorisation limits. These still function, but only against the values they were written for. They do not constrain aggregate exposure across many small automated actions, and they say nothing about frequency. Audit evidence. The evidence of a human decision is the decision-maker, who can be interviewed. The evidence of a model's decision is an output. You cannot ask it what it was thinking and get an answer you should rely on.
The five controls that replace them
This is the part worth designing carefully, because it is genuinely different rather than just more. Preventive boundaries. Hard limits the system cannot exceed regardless of what it concludes — vendor allowlists, value ceilings, account restrictions, frequency caps. These are stronger than any detective control because they do not depend on anyone noticing. Statistical review. Stop checking every item and start checking a sample properly, weighted toward the unusual. A hundred items reviewed carelessly is worth less than ten reviewed properly. Behavioural monitoring. Track the distribution of the system's decisions over time. A shift in the exception rate, the average value, or the vendor mix is the earliest available signal that something has changed upstream. Independent reconciliation. A deterministic check, outside the automated process, that the aggregate is right. Total posted equals total invoiced. Simple, boring, and the control auditors will find most convincing. Change control on the configuration. The prompt, the threshold, the tool list and the model version are all control-relevant. In most organisations they can currently be changed by one person with no record, which would be unthinkable for a general ledger setting.
What to give your auditor
A short document: what the process does, which controls are preventive and which detective, where the independent reconciliation sits, how the sample is drawn and sized, who owns configuration changes and how they are recorded. Produce it before the fieldwork. An auditor encountering an automated process with no control narrative expands scope, and that costs more than writing the document.
Practical Guidance for Controls Design Workshop
- Stop counting two automated steps as segregation of duties.
- Replace item-by-item approval with properly sized sampling.
- Add aggregate and frequency limits, not just value limits.
- Build one deterministic reconciliation outside the automation.
- Put prompts, thresholds and model versions under change control.
- Monitor decision distributions as a leading indicator.
- Weight samples toward the unusual, not evenly.
- Write the control narrative before the auditor asks.
The Regional Angle
The first practical issue is one of scale. A great many regional finance functions are small enough that segregation of duties was already partly notional — in a four-person team the same person often prepares and posts, with the control supplied by the finance manager's review or the owner's sight of the payment run. Automation is frequently pitched here as the solution to that weakness, and it can be, but only if the independence is designed in deliberately. The pattern that works in small teams is a preventive boundary plus a deterministic reconciliation, because both function without a second person. Relying on a second automated step to supply the independence a second person would have supplied leaves you with fewer real controls than you started with. The second concerns the audit environment you will be explaining this to. Regional external audit practice varies considerably between the international firms' local offices and smaller local practices, and the smaller practices in particular are encountering model-operated processes for the first time. The practical consequence is not scepticism but substitution: an auditor who cannot evaluate the automated control reverts to substantive testing, which means more sampling, more document requests and a longer fieldwork period at your cost. The control narrative is therefore worth more here than in a market with established guidance, and the most useful single item in it is the deterministic reconciliation, because it is a control any auditor can test without understanding the model at all. The third is about statutory processes where the consequence of an error is external rather than internal. E-invoicing submissions to the Saudi tax authority, value added tax returns across several Gulf jurisdictions, wage protection files and end-of-service calculations all produce filings that are difficult to correct and visible to a regulator. These deserve a different control posture from internal postings: keep a human authorisation at the submission point, run the deterministic reconciliation before rather than after, and do not let the automation own the filing calendar. The internal ledger tolerates a correcting entry. A cleared invoice under a tax authority's regime does not, and a group operating across four jurisdictions has four separate versions of that problem.
The objection worth taking seriously
The strongest objection is that this overstates the novelty. Automated controls have existed in enterprise systems for thirty years — three-way match tolerances, system-enforced approval hierarchies, automatic posting rules — and auditors have tested them successfully throughout by examining the configuration and testing a sample of outputs. The existing framework already accommodates machine-performed controls perfectly well, and declaring that segregation of duties has collapsed because the automation is now statistical rather than deterministic is a consulting narrative in search of a workshop fee. The historical point is correct, and the testing approach — examine the configuration, test the outputs — does largely carry over. The difficulty is with the first half of that approach. Testing a configuration works because a deterministic rule's configuration fully determines its behaviour: read the tolerance, and you know what the system will do in every case. A model's configuration does not have that property. You can read the prompt, the threshold and the tool list and still not know how it will handle a case nobody anticipated, which means configuration testing gives assurance about intent rather than behaviour. That is precisely why the weight has to shift toward sampling of actual outputs, behavioural monitoring, and preventive boundaries that constrain outcomes rather than describing intentions. The framework survives. The mix of controls inside it has to change, and organisations that leave the mix unchanged will hold a control environment that documents diligence rather than producing it.
Common Questions
Can a second agent provide independent review?
Only if it is genuinely independent — different vendor, different model, different data path — and even then it is weaker than a deterministic reconciliation. Two instances of the same system are one control.
How large should the sample be?
Large enough to detect the error rate you care about, weighted toward unusual transactions, and reviewed properly. A smaller sample examined carefully beats a larger one signed off in bulk.
Do we need to log the model's reasoning?
Log the inputs, the decision and the limit that permitted it. Retained explanations are useful context and are not evidence of why the system actually behaved as it did.
What should we expect over the next twelve months?
Expect the audit firms and standard setters to publish guidance on controls over model-operated processes during this year. Expect early findings to concentrate on configuration change control, because almost nobody has it. Expect preventive boundaries to become the default recommendation over detective review. And expect the organisations with a written control narrative to have materially shorter audits than those without.
Controls Design Workshop — we rebuild the control set around what the automation actually does, and write the narrative your auditor will ask for.
