Back Office / Source date:

Audit Trails for AI-Assisted Finance Decisions

Auditors require traceable inputs and rationale, which most AI tooling did not capture by default.

Illustration of a retained finance evidence pack linking inputs, review notes and the approved decision.

Audit planning for the current financial year is underway, and a new question has started appearing in the pre-planning questionnaires: describe where artificial intelligence is used in the preparation of the financial statements, and the controls over it. Most finance teams cannot answer either half. Not because they are hiding anything, but because the tools arrived through operations rather than through a controls review, and nobody wrote down what they touch.

A log records what the system did. Evidence explains why a person was entitled to rely on it. Almost every deployment has the first and none of the second

That distinction is the whole of the work, and it is why exporting an application log in November will not answer the question.

What an auditor is actually testing

Not the model. This is the most common misunderstanding in the current debate, and it leads finance teams into pointless arguments about explainability. An auditor tests a control: was it designed to catch the misstatement it was meant to catch, did it operate consistently throughout the period, and can you demonstrate both. Those three questions are unchanged by artificial intelligence. What changes is the meaning of "operated" — because the thing performing the step is no longer a person whose initials are on the document. So the correct response to the questionnaire is not a technical defence of the model. It is a control description with evidence attached.

Five things a log almost never captures

The configuration in force at the time. Which model version, which prompt, which settings. This changes without announcement in most hosted products. The complete input. Not just the document, but any retrieved context that influenced the output. An answer assembled from three internal sources is not reproducible from the question alone. What the reviewer actually saw. If a person approved, the evidence needs to show the screen they approved on, not merely that a record was updated. The rejected alternatives. Where the system selected among candidate matches or codings, the discarded options are frequently the most informative part of the record. An immutable timestamp. A log that the team operating the system can edit is a working paper, not evidence.

Retain evidence of the decision, not only activityQualitative evidence requirements proposed in the article. Retention periods and audit sufficiency depend on the applicable regime and engagement.
Evidence elementWhat it helps establish
Configuration identifierModel, prompt and settings in force
Complete input and retrieved contextInformation that influenced the output
Reviewer presentationWhat the approving person saw
Rejected alternativesCandidate choices behind the decision
Protected timestamp and change recordWhen the step occurred and what later changed

Qualitative summary of this article's source text, not a measured outcome or performance estimate.

Three control patterns that hold up

Reperformance sampling. A defined monthly sample, independently re-done by a person, with the comparison retained. Unglamorous, entirely familiar to auditors, and the fastest route to a testable control. The sample size argument is easier than the explainability argument. Authorisation at a threshold. Above a stated value, or outside a stated confidence band, a named person authorises and their basis is captured. This converts an opaque automated step into a documented human decision at exactly the points where the money is. Completeness reconciliation. Prove that the population the system processed reconciles to the population in the ledger. This catches the failure mode nobody logs: not a wrong answer, but a silently missing item. In practice it finds more problems than accuracy testing does.

Reproducibility is the uncomfortable part

If a hosted model is upgraded in August, the March output cannot be reproduced. That is a genuine break with how finance evidence has always worked, and pretending otherwise in an audit meeting damages credibility. The workable position is to stop promising reproduction and start retaining artefacts. Store the inputs, the output and the configuration identifier at the moment of use, so the record is complete even though the computation cannot be re-run. Then put a notification clause about material model changes into the contract, and record the change dates in the same place you record any other system change. An auditor can work with "here is what it produced and here is what it was" — they cannot work with nothing.

Tier the effort by consequence

Not everything needs the same standard. Coding a low-value expense and estimating an impairment are not the same risk, and applying one evidence policy to both guarantees you over-engineer the first and under-engineer the second. Sort AI-touched steps by financial statement impact and by reversibility, then set the evidence requirement per tier. Auditors respond well to a documented tiering rationale; they respond badly to a uniform policy that visibly nobody follows.

What to hand over

Five artefacts, prepared once and maintained: an inventory of where AI is used in the finance process; a matrix mapping each use to its control; evidence samples from the period; a change log covering configuration and vendor updates; and the exception record showing what escalated and how it was resolved. Teams that walk into planning with those five have a short conversation. Teams that do not spend six weeks assembling them under time pressure in the middle of the close.

Practical Guidance for AI Controls Assessment

  • Inventory every AI-touched step in the financial reporting process.
  • Describe controls, not models, in audit correspondence.
  • Capture configuration identifiers with every logged output.
  • Retain inputs and outputs rather than promising reproduction.
  • Run monthly reperformance sampling and keep the comparison.
  • Reconcile populations to catch silent omissions.
  • Make timestamps immutable and outside the operating team's control.
  • Tier evidence by financial impact, and document the rationale.

The Regional Angle

The first practical trap is retention, and it is a statutory question rather than an IT policy choice. Tax legislation across the region sets minimum record-keeping periods measured in years — five for value added tax in the Emirates, longer for real estate records, seven under the corporate tax regime, with comparable obligations in Saudi Arabia reinforced by e-invoicing archiving requirements. Meanwhile the default log retention in most software-as-a-service products is somewhere between thirty and ninety days, and the audit trail inside a hosted AI feature is frequently shorter still. An organisation can therefore be fully compliant in its accounting system and non-compliant in the evidence that explains how a number in that system was derived. Check the actual retention setting of every tool touching a taxable record, export the artefacts to your own storage on a schedule if the vendor cannot hold them for the statutory period, and record the retention basis alongside the inventory. This is the cheapest finding to avoid and one of the most commonly missed. The second is a structural feature of regional groups: it is normal to have several audit firms across one group. Free zone authorities maintain approved auditor lists, a joint venture partner insists on their own firm, one entity keeps a long-standing local relationship, and the group auditor covers the consolidation. A single AI-assisted control operating across shared services therefore has to be evidenced separately to three or four different firms, each with its own questionnaire, sampling preference and view of what constitutes sufficient documentation. Build the evidence pack at entity granularity from the outset, even where the control is centralised — samples tagged by entity, exceptions attributable to an entity, and inventories that can be filtered. Producing a group-level pack and then disaggregating it under deadline is the version of this that consumes an entire finance team's January. The third is that the stakes attached to the audit opinion are higher here than the reporting obligation alone suggests. Audited statements are not merely a shareholder document: they are routinely required for trade licence renewal, demanded by banks for facility and overdraft renewals, and now feed directly into the corporate tax position. A qualification or a significant control deficiency does not stay inside the audit report; it surfaces in a facility review at the least convenient moment. That asymmetry argues for spending the evidence effort early, on the small number of AI-touched steps with real financial statement impact, rather than distributing it thinly across every tool anybody has installed.

The objection worth taking seriously

The strongest objection is that this is retrospective paperwork wrapped around a technology that is still finding its footing. No auditor has asked most finance teams anything about artificial intelligence yet. The standards bodies have not issued anything definitive. Building a control and evidence framework now means designing against a requirement that does not exist, and the likely outcome is an elaborate apparatus that gets rewritten the moment guidance actually arrives. The factual part of that is right. The questionnaires appearing this spring are early and inconsistent, and anyone building a comprehensive framework today is guessing at the details. The reason to act anyway is timing rather than content. Evidence of a control operating is inherently contemporaneous — it cannot be created afterwards. When the requirement does firm up, the organisations that started logging configuration identifiers and running a monthly sample will have a year of operating history, and the ones that waited will have a policy document with no evidence behind it and a control that, in audit terms, has been in place for six weeks. The incremental cost right now is small: three or four extra fields in a log, a sampling routine that a junior can run in an afternoon each month, and a spreadsheet inventory. None of that is a framework. It is the minimum that makes a future framework cheap instead of painful, and it is useful immediately for entirely non-audit reasons — the same artefacts are what you need when a supplier disputes a coding, or when somebody asks how a payment was approved.

Common Questions

Do we need to explain how the model works?

No. You need to demonstrate that a control over its output operated. Explainability is a useful property, not an audit requirement.

Is a human approval click sufficient evidence?

Only if you can show what was presented at the moment of approval. A click on a screen nobody can reconstruct is weak evidence of review.

What if the vendor will not tell us when the model changes?

Treat it as a control weakness, document it, and compensate with heavier sampling. It is also a reasonable thing to raise at renewal.

What should we expect over the next twelve months?

Expect audit firms to standardise their questionnaires during the year, and expect the questions to get more specific rather than less. Expect at least one professional body to publish guidance on evidence for AI-assisted processes, and expect it to say roughly what good practice already says: control the output, not the model. Expect vendors to start marketing audit trail features once buyers begin failing them in procurement. And expect the first uncomfortable conversations to arrive not from auditors but from tax authorities asking how a figure in a filed return was arrived at.


AI Controls Assessment — we inventory where AI touches your numbers, design the controls your auditors will test, and make sure the evidence exists before anyone asks for it.

Continue reading

Talk to OPS

Start with the operating problem.