The phrase AI copilot is being attached to enterprise software at a speed that should make anyone suspicious, and the first thing worth saying is that the word is not new. SAP shipped a product literally called CoPilot several years ago — a digital assistant bolted onto its business suite — and almost nobody used it. Oracle has offered a digital assistant for its cloud applications since 2018. Salesforce has been selling Einstein for more than five years. What has changed is not the vocabulary. It is that a coding assistant in technical preview since last summer has visibly altered how a large number of developers spend their day, and that Microsoft has been quietly using a language model inside Power Apps since last May to turn a sentence into a working formula. For the first time the pattern — the software drafts, the human accepts — is producing measurable behaviour change rather than demo applause. That pattern is now heading for the systems that run finance, supply chain and procurement, and the question for anyone who owns one is what has to be true before a suggestion can be allowed near a ledger.
Three generations, only one of them new
It helps to separate what is actually arriving from what has been here for years. The first generation runs in the background and predicts: lead scoring, likely payment dates, demand forecasts, anomaly flags on expense claims. It has existed in mainstream suites since around 2016, and in most installations it is either switched off or switched on and ignored, because acting on a score required a workflow nobody built. The second generation extracts: reading an invoice or a delivery note and turning it into fields. Narrow, unglamorous, and — finally — good enough to deploy, which is a subject in its own right. The third generation is the interesting one. It drafts the thing the user was about to do, in the place they were about to do it, and waits to be accepted or rejected. A human is in the loop by design, which sounds reassuring and is actually where all the difficulty lives.
Five jobs a copilot would earn its keep doing
Strip away the marketing and the credible near-term uses in an enterprise system share one shape: a small decision, made thousands of times, with a long history of correct answers already sitting in your database. Coding a supplier invoice to the right account, cost centre and tax code based on what happened to the last two hundred invoices from that supplier. Matching an incoming remittance against open invoices when the reference is missing or wrong. Turning an email request into a draft purchase requisition with the right item codes. Explaining a variance — assembling the query that shows why a cost centre moved, rather than making the user learn the reporting tool. And writing the filter, formula or report definition that a competent finance person can describe in a sentence but cannot express in the software's own syntax. Notice what is absent from that list: deciding anything material. The good uses are clerical.
The accept rate, and why it is dangerous
There is only one metric that tells you whether an assistant is working: the proportion of suggestions accepted without modification, measured against a sample of independently checked outcomes. And there is a trap sitting inside it. An assistant that is right eighty-five per cent of the time is good enough to be useful and good enough to train its users to stop reading. Acceptance stops being a judgement and becomes a keystroke. That is tolerable when the suggestion is a report filter and unacceptable when it is a posting, a payment run or a credit limit, so the design question is not how accurate the model is — it is which suggestions you allow to be accepted cheaply and which must cost the user something to confirm.
What your system needs before suggestions are safe
Five things, and most enterprise systems currently have two of them. A genuine proposed state, distinct from draft and distinct from posted, so a suggestion can exist in the system without being an entry. An accountability field. Every record needs to carry who or what proposed it and who accepted it, as separate values. Today your audit trail has one field — created by — and it will contain the name of a human being who clicked a button. When someone asks in two years how a misclassification happened, the difference between typed it and accepted it is the whole answer, and you cannot reconstruct it later. Confidence, surfaced and thresholded. A suggestion shown with no indication of certainty invites uniform trust. Below a threshold, show nothing; silence is a better behaviour than a plausible guess. Reversibility with a clean trail. Accepted suggestions will be wrong. Reversal must be a normal operation that leaves evidence, not a journal entry someone writes to compensate. Sampling review. Somebody checks a random sample of accepted suggestions every week, permanently. Acceptance is not evidence of correctness, and the only way to know your accept rate is honest is to test it against reality on a schedule.
Keep a proposed state
Hold the suggestion apart from posted entries.
Record separate identities
Store who proposed the entry and who accepted it.
Apply consequence thresholds
Suppress weak suggestions and require stronger confirmation for costly mistakes.
Preserve reversal evidence
Make reversal a normal operation with a clean trail.
Sample accepted work
Check outcomes independently; acceptance alone is not correctness.
Qualitative summary of this article's source text, not a measured outcome or performance estimate.
Your history is the model
A suggestion engine trained on your transactions will reproduce your transactions, including the parts your auditor keeps asking about. If travel has been coded to three different accounts depending on who processed it, the assistant will learn the distribution rather than the rule. If one entity has been putting freight into cost of sales and another into overheads, pooling their history produces a model that is confidently wrong in both. So the prerequisites are dull and familiar: consistent master data, a coding policy that is actually followed, and enough clean history for the patterns to be real. There is also a segregation-of-duties question worth asking early. If the same assistant can propose a supplier record and propose a payment, you have introduced a route around a control that took years to establish.
Practical Guidance for AI-Ready ERP Assessment
- Do not buy an AI module this year. Test whatever your existing vendor already ships, on your own data.
- Add the accountability field now — proposed by and accepted by, stored separately — because you will need it either way.
- Classify suggestion types by consequence, and require real confirmation only where reversal is expensive.
- Insist on accept and reject telemetry as a procurement requirement, not a roadmap item.
- Start with invoice coding and remittance matching, where a large history of correct answers already exists.
- Fix the coding policy before training anything on it, or you will automate your inconsistencies.
- Run a weekly sample review of accepted suggestions and publish the error rate internally.
- Ask where inference happens and what leaves your tenant before enabling any assistant feature.
The Regional Angle
Three regional realities change the answer here, and the first is a straightforward quality question that vendor demonstrations are designed to avoid. Language model performance in Arabic remains well behind English, and regional business documents are rarely in one language. A supplier invoice carries an Arabic trade name and an English description. A delivery note is handwritten. A narration field contains transliterated Arabic typed on an English keyboard, which no tokeniser handles gracefully. The practical consequence is that a feature which performs well in an English-only workflow can degrade sharply on your actual document set, and the degradation will not be uniform — it will be concentrated in exactly the suppliers and cost centres where manual coding was already hardest. The test is not the vendor's sample set. It is last quarter's real documents, scored by someone who knows what the right answer was, field by field. The second is that many regional businesses simply do not have much usable history. A large cohort re-implemented or substantially reconfigured around the introduction of value added tax four years ago; others migrated during the last two years. Three or four years of postings in a single mid-sized entity is a thin basis for learning, and the obvious remedy — pooling history across the group's entities — fails where each entity was configured by a different partner with a different chart of accounts. Before evaluating any assistant, work out honestly how many correctly coded examples exist for the decision you want automated. If the answer is a few thousand spread across four incompatible conventions, the constraint is not the model. The third is newly legal rather than technical. Almost every assistant feature sends content somewhere to be processed, and that content is your invoice narrations, supplier names, contract text and occasionally employee data. The Emirati federal personal data protection law came into force at the start of this month, Saudi Arabia's law takes effect in March, and both put weight on knowing where personal data is processed and on what basis it leaves the country. Enabling a suggestion feature with a tick box can create a cross-border transfer that appears in no register, has no processing agreement behind it, and was authorised by whoever ran the pilot. Add one question to the assessment template — where does the inference call execute, and what is retained — and put the answer in your records of processing before the feature is switched on rather than after a regulator asks.
The objection worth taking seriously
The strongest objection is that we have been here repeatedly since 2016, and the pattern is always the same: a vendor announces intelligence in the suite, a handful of features appear behind an additional licence, nobody configures the workflow required to act on them, and three years later the feature is quietly deprecated. SAP's own assistant, which carried exactly this name, is the best available evidence. On this reading the copilot framing is a rename of autocomplete, and the real constraint has never been the model — it is master data, process discipline and the fact that nobody owns the decision the software is offering to help with. No assistant fixes any of those. That history is accurate and the caution is earned. Two things are nevertheless different this time. The interface is natural language rather than structured configuration, which removes the implementation project that killed the previous generation — the earlier features needed a data scientist and a workflow designer before they produced anything, and these do not. And the acceptance loop generates measurement for free: for the first time you can see, per user and per transaction type, whether the software's suggestion was any good. Background prediction never produced that evidence, which is precisely why nobody could defend keeping it switched on. The correct posture is therefore neither enthusiasm nor dismissal. It is to build the plumbing — proposed state, accountability fields, sampling review — which is cheap, useful on its own terms, and the only thing that makes the next two years of vendor announcements evaluable rather than merely irritating.
Common Questions
Should we wait for our vendor or buy a specialist tool?
Wait for the vendor for anything touching the ledger, because integration and audit trail matter more than model quality. Consider specialists only for narrow, bounded tasks like document extraction.
Who is accountable when an accepted suggestion is wrong?
The person who accepted it, which is exactly why they must be able to see why it was suggested and must not be measured on throughput alone.
Does this reduce headcount?
Not at the scale being implied. It compresses the clerical middle of a process while adding review work, and the net effect in the first year is usually a redistribution of effort rather than a reduction.
What should we expect over the next twelve months?
Expect assistants to be announced at every major vendor event this year, priced as per-user add-ons rather than included. Expect the coding assistant currently in preview to reach general availability and to make the term unavoidable. Expect the first serious arguments about liability for an accepted suggestion, and the first audit finding that turns on whether an entry was typed or accepted. And expect the early production wins to be narrower and duller than the announcements — invoice coding, matching, report writing — which is the strongest available signal that the underlying capability is real.
AI-Ready ERP Assessment — we test suggestion features against your own documents, count the usable history you actually have, and put the accountability trail in place before anyone clicks accept.
