Three weeks after a free conversational AI tool arrived, finance and shared services teams are asking a different question from the rest of the organisation. Marketing wants to know whether it writes good copy. The back office wants to know whether anything it produces can survive an audit. That is the right question, and the answer is more useful than either the enthusiasts or the sceptics suggest, provided one distinction is held firmly.
It can draft the memo explaining the accrual. It cannot be the reason the accrual is that amount
Everything sensible about generative AI in a back office follows from that line. These systems are strong at turning unstructured input into structured output, at explaining complicated text in plain language, and at producing first drafts of narrative. They are weak at arithmetic, unreliable about facts they were not given, and structurally incapable of being an authority. They also produce different answers to the same question on different days, which is disqualifying anywhere reproducibility is part of the control. So the dividing line is not seniority or risk appetite. It is whether the output determines a number, a classification with a tax or accounting consequence, an approval, or a filing. On that side of the line, the tool contributes nothing you can defend. On the other side sits a surprisingly large amount of back-office work.
Five places it already earns its keep
Narrative that accompanies numbers. Variance commentary, board memos, audit responses, policy and procedure documents. The numbers come from the system; the prose comes faster than it used to. Explanation and translation. Turning a tax circular, a contract clause, a bank letter or an obscure system error into language a clerk can act on. This is the highest-value use in most shared service centres and the least discussed. Unstructured input into structured candidates. Remittance advice buried in an email body, a supplier's chaotic statement, a expense narrative that needs coding. Note that dedicated document processing tools remain better for structured documents; the language model earns its place on the messy long tail those tools reject. Writing the query rather than the answer. A spreadsheet formula, a database query, a reconciliation script. The output is testable, which converts an unreliable assistant into a useful one. First-pass triage. Classifying incoming requests, routing correspondence, drafting the standard reply for a person to send. Cheap to check, easy to reverse.
| Assistance candidate | Verification that stays outside the draft |
|---|---|
| Variance commentary | Reconcile figures to the source system. |
| Plain-language explanation | Check requirements against the original text. |
| Query or formula draft | Test the code and reconcile its outputs. |
| Correspondence triage | Keep accountable human review before consequential action. |
Qualitative summary of this article's source text, not a measured outcome or performance estimate.
The control question nobody has asked yet
Segregation of duties assumes a preparer and a reviewer who are different people. If a model prepares two hundred journal narratives and one accountant approves them all, you do not have a strengthened control. You have a reviewer with no preparer, reviewing at a volume that guarantees the review is nominal. This is the part that will eventually interest your auditors, and it has an unglamorous answer: keep a named human preparer for anything that enters the ledger, record where assistance was used, and set review sampling based on what a person can genuinely check rather than on what the queue contains. Write that down before somebody asks.
The economics changed in one direction only
Automating a back-office exception used to require a rules engine or a model, which meant an engineering project, which meant only high-volume processes justified the effort. A large share of that work is now a prompt, so the volume threshold for automating a narrow process has collapsed. That is genuinely new. What has not changed is verification cost. Somebody still reads the output, and for anything consequential that person must be competent enough to catch a confident error. Per item, the total cost may barely move. The gain shows up as cycle time and as capacity redirected from typing to checking — which is why measuring this in headcount will make it look like a failure, and measuring it in days-to-close and rework rate will make it look like what it is.
Ninety days, run properly
Pick two processes with high text volume and low financial consequence: supplier correspondence and procedure documentation are the usual candidates, with audit request responses a close third. Baseline the current cycle time and rework rate before you start, because nobody remembers what normal looked like afterwards. Give each one a named owner. Run for a quarter. Decide on the numbers. And resolve the data route first. Back-office data is the worst possible material to paste into a consumer service — payroll files, customer ledgers, supplier bank details, unpublished results. Either use synthetic or redacted material for the pilot, or make an enterprise arrangement with contractual terms on training and retention a procurement item for the first quarter. There is no defensible middle option, and the middle option is what most teams are currently doing.
Practical Guidance for GenAI Back Office Strategy
- Draw the ledger line explicitly: drafting and explanation on one side, amounts and approvals on the other.
- Keep a named human preparer for anything entering the financial records.
- Pilot two low-consequence, text-heavy processes with a baseline measured first.
- Measure cycle time and rework, not headcount reduction.
- Resolve the data route before the first real document is used.
- Use it to write queries and formulas, where the output can be tested.
- Validate generated procedure documents against a system trace, never against memory.
- Document where assistance was used, because your auditor will eventually ask.
The Regional Angle
Three regional factors make this both more useful and more dangerous here than the general discussion suggests. The first is the sheer weight of regulatory text landing on regional finance teams right now. The Saudi e-invoicing integration phase begins in two weeks, and the technical documentation surrounding it — specifications, validation rules, implementation resolutions, phased waves — is exactly the kind of dense material that a language model summarises beautifully and occasionally invents. Used well, it turns a forty-page circular into the three questions you need to ask your tax adviser, which is genuine value in a team that has no time to read it. Used badly, somebody treats the summary as the requirement and configures a system against a rule that does not exist. The discipline is simple and should be stated as policy: the model drafts the question, the adviser gives the answer, and nothing derived from a summary reaches a configuration screen without a citation to the actual source text. The second is the documentation burden arriving with corporate tax in the Emirates for financial years beginning from the middle of next year. Groups here are about to produce categories of document they have never produced — tax positions, intercompany policies, supporting memoranda, transfer pricing documentation — in volumes that will overwhelm small finance teams. This is the largest new drafting workload the regional back office has faced in years, and it is precisely where assistance helps with structure and language. It is also precisely where a fabricated reference to a non-existent provision is fatal, because the reader is a tax authority. Let the tool produce the skeleton and the plain-language sections; let a professional supply every fact, every figure and every legal reference. The third is about how process knowledge is actually held in this region. Regional back offices run on visa-sponsored staff with meaningful turnover, and a great deal of process knowledge lives in individual habit rather than in documentation. When the person who knew how the intercompany reconciliation worked leaves, the temptation to have a model write the procedure document is enormous — and it will produce a confident, plausible, well-structured description of a process that was never followed. Generated documentation must be validated against a system trace: run the transaction, screenshot the steps, reconcile the document to what the system actually does. Written that way it is a genuine gain over the folder of undated Word files most groups currently rely on. Written from the model's imagination, it becomes the control evidence you hand to an auditor, and it is fiction.
The objection worth taking seriously
The strongest objection is that the back office is the last place on earth for a probabilistic tool. Finance runs on determinism, reproducibility and an audit trail. A system that answers differently on Tuesday than it did on Monday cannot sit inside a control environment, auditors will not accept it as evidence, the productivity claims circulating this month are anecdotal at best, and the responsible course is to wait eighteen months until the enterprise software vendors embed this properly, with permissions, logging and reproducibility built in. That argument is correct for everything that touches the ledger, and the wait-for-the-vendor case is genuinely stronger in the back office than anywhere else in the business, because controls and audit trails are not features you can bolt on afterwards. But the drafting and explanation layer sits outside the ledger and does not need determinism — a variance commentary is not a control, and a plain-language explanation of a tax circular is not evidence. Waiting also is not neutral: your team is using this today, on personal accounts, with real documents, and the eighteen-month wait is eighteen months of that. And the capability worth building in the meantime is not prompting. It is the organisational habit of specifying work precisely and verifying output rigorously, which is the same habit you will need when the vendor finally ships something embedded.
Common Questions
Can it do the reconciliation?
No. It can explain a reconciliation, draft the commentary on the variances, and write the query that produces the matching. The matching itself belongs to a deterministic system.
Will auditors accept work produced this way?
For narrative and documentation, with a human preparer and a clear trail, this is unlikely to trouble anyone. For anything constituting evidence or determining an amount, assume it will be challenged and plan accordingly.
Does this reduce headcount?
Not in the first year, and pitching it that way will produce a pilot that quietly fails. It moves effort from producing text to checking it, and the honest business case is cycle time and capacity.
What should we expect over the next twelve months?
Expect enterprise software and finance vendors to announce embedded assistants during the year and to ship them later than announced, with the first credible versions confined to narrative and search rather than to transactions. Expect enterprise terms carrying no-training and retention commitments to become available early in the year, which is what unblocks serious back-office use. Expect the audit profession to publish guidance, and expect the first questions about AI use in preparing financial information to arrive during next year's audit cycle. Expect the measurable wins to be in documentation, correspondence and explanation rather than in processing. And expect at least one well-publicised case of a wrong number traced back to a model-drafted memo that nobody checked — which will do more to establish sensible practice than any policy written this quarter.
GenAI Back Office Strategy — we draw the line between drafting and deciding, pilot the two processes where it actually pays, and make sure your controls still hold when the preparer is a machine.
