Budget season opens with a question that did not exist eighteen months ago. The assistant licences bought during last year's enthusiasm come up for renewal, the seat count is material, and somebody on the finance committee wants to know what they bought. At most organisations the honest answer is that nobody knows. There is a usage dashboard, a set of enthusiastic quotes, and a survey in which people estimated how much time they saved.
A productivity gain you cannot see in a cycle time, a headcount or a backlog is a productivity gain you did not have
That is a harsh test and it is the right one. Value that never reaches a queue, a deadline or a cost line is value that exists only in the deck.
Four ways to measure this, in ascending order of honesty
Ask people how much time they saved. The most common method and the least reliable. It measures perception, it anchors on whatever number the question suggests, and staff who like the tool will report savings whether or not their output changed. Useful for sentiment. Worthless for arithmetic. Read the vendor telemetry. Active users, prompts per week, acceptance rates. This tells you whether the tool is being used, which is necessary and nowhere near sufficient. High engagement with a tool that changes nothing downstream is the most expensive result available. Compare the team before and after. Intuitive and badly confounded. Headcount moved, the process changed, the quarter was busier, and — unavoidably — you announced a productivity initiative to people whose output you were about to examine. Stagger the rollout and keep a holdout. Give the tool to half the population now and half in twelve weeks, then compare. This is the only method that answers the question, and it costs nothing extra because you were going to roll out in waves anyway. The objection is always that withholding a useful tool is unfair. You are not withholding it. You are sequencing it, which is what every rollout does. The only change is deciding the order deliberately and measuring while you wait.
Use metrics that already exist
The strongest signal that a measurement programme will be abandoned is that someone had to invent the metric for it. Pick things the business already tracks and already cares about: cycle time for a defined process, age of the oldest item in a backlog, rework or error rate, throughput per person, cost per transaction, days to close. If a copilot is worth what the vendor claims, at least one of those numbers moves within a quarter. If none of them move, the value is real to the individual and invisible to the organisation — which is a legitimate finding and should change what you buy next.
Where the gains actually appear in an ERP context
Three areas behave very differently, and lumping them together is why blended figures are meaningless. Reporting and query work is where the clearest gains sit. The measurable outcome is the analyst request backlog and the time from question asked to answer received. Configuration and change work benefits substantially — drafting, documentation, test scripts, explaining why something behaves as it does. Measure change request lead time. Transaction processing is where the smallest gains are, and where boards expect the largest. The constraint on an invoice is rarely the typing; it is the approval waiting in somebody's inbox and the master data that does not match. Assistants do not fix either.
| Use case | Suggested outcome | Keep separate |
|---|---|---|
| Reporting and queries | Request backlog and time to answer | Prompts and active users |
| Configuration and changes | Change-request lead time | Drafting activity alone |
| Transaction processing | Cycle time, rework and cost per transaction | Approval delays and master-data cleanup |
Qualitative summary of this article's source text, not a measured outcome or performance estimate.
The attribution trap
Most first-year AI gains are process cleanup. To deploy the tool you documented the process, fixed the master data, removed three approval steps and retired a spreadsheet. Output improved. Some of that is the assistant and a great deal of it is the tidying. That is genuine value and you should take it. But label it correctly, because if you attribute it all to the licence you will buy more licences next year expecting the same curve, and you will not get it. The cleanup was a one-off.
What the board should receive
Not a percentage. Three things: an adoption number as the leading indicator, one outcome metric per use case as the lagging indicator, and fully loaded cost per active user per month. Then a stop rule — stated in advance, in writing: if the outcome metric has not moved by this date, we stop at this seat count. Nothing else you do will buy as much credibility with a finance committee as naming the conditions under which you would walk away, before you know the answer.
Practical Guidance for AI Value Measurement Review
- Sequence the rollout so a holdout group exists by default.
- Choose metrics the business already tracks, never invented ones.
- Separate reporting, configuration and transaction use cases.
- Treat vendor telemetry as leading, never as outcome.
- Baseline before switching anything on, for at least six weeks.
- Attribute the process cleanup separately from the tool.
- Report cost per active user, not cost per licence.
- Publish a stop rule before the pilot begins.
The Regional Angle
Three things change how this plays for groups operating here, and the first is that the baseline does not exist. In a great many regional finance and operations functions, nobody has ever measured cycle time. There is no record of how long an invoice actually takes from receipt to posting, how long a change request waits, or how old the oldest open item is — because the function has always been judged on whether the close happened and whether the auditors were satisfied. When you start an AI measurement programme, the first six weeks are not measuring the AI. They are the first time anyone has measured the process, and that dataset is frequently worth more than the tool being evaluated. Treat the baseline as a deliverable in its own right, resource it properly, and expect at least one uncomfortable discovery that has nothing to do with artificial intelligence. Teams that skip this step end up comparing a measured period against a remembered one, which always favours the new thing. The second is that value varies enormously by entity within the same group, and group-level measurement hides it completely. A drafting assistant is transformative in the entity where three accountants support four hundred million in revenue and produce commentary in English for a regional head office. It is close to worthless in the sister entity where the finance manager writes a short summary in Arabic for a family shareholder who will ask questions verbally. Same group, same licence, same use case, entirely different value. Measure per entity and per process, choose the pilot population by where the backlog is worst rather than by which managing director is most enthusiastic, and be prepared to deploy at very different densities across the group. A uniform seat allocation across entities is the most common way regional groups overspend on this. The third is political, and it decides whether any of the above survives contact with reality. In owner-led and family-controlled organisations — which is most of the regional private sector — once the chairman has decided that AI is a priority, producing a negative measurement becomes a career question rather than an analytical one. Nobody running the pilot will report that it did not work. The structural fix is to separate the measuring from the deploying: agree the metric and the stop rule with the sponsor in writing before anything is switched on, and have the result compiled and presented by finance or internal audit rather than by the team that championed the tool. If that separation feels unnecessary, that is usually the strongest evidence it is needed.
The objection worth taking seriously
The strongest objection is that this is disproportionate. An assistant licence costs a few tens of dollars per user per month, which is trivial against a loaded salary. Demanding a staggered rollout, a holdout group and a controlled comparison before spending that is exactly the kind of governance that makes organisations slow, and the measurement programme can easily cost more than the software it is measuring. Meanwhile competitors are simply deploying, staff want the tools, and in a labour market this tight, refusing to give people modern software carries a retention cost that no spreadsheet will capture. That is right at small scale, and the proportionality point is correct. For fifty seats, deploy and move on. Do not build a research programme to validate a rounding error. The measurement matters at the next decision, not this one. The question that needs evidence is not the fifty-seat pilot; it is the two-thousand-seat rollout, the multi-year enterprise agreement, and the headcount plan that gets built on an assumed efficiency gain. Those are seven-figure commitments, and they are routinely justified with a survey. The version recommended here is also far cheaper than it sounds: sequence a rollout you were already sequencing, track three numbers the business already produces, and write down in advance what would make you stop. That is a fortnight of someone's attention, and it is the difference between knowing what you bought and believing what you were told.
Common Questions
How long before we should expect to see movement?
One quarter for reporting and drafting use cases. Two or more for anything touching a process with approvals in it, because the constraint is not the part the assistant touches.
Is acceptance rate a useful metric?
As a leading indicator of whether the tool is fit for the work, yes. As evidence of value, no — accepting a suggestion is not the same as producing more.
What if the measurement shows no gain?
That is a result, not a failure. It usually means the use case was wrong rather than the technology, and it saves you from scaling the wrong deployment.
What should we expect over the next twelve months?
Expect the vendors to publish their own return-on-investment studies during the year; read them as marketing, because they will be. Expect assistant pricing to be folded into suite and edition pricing, which will make cost attribution harder rather than easier — so capture your current per-seat cost now, while it is still visible as a line. Expect the first genuinely independent field studies to appear during the year and to report gains that are real, narrower and more task-specific than the current claims. And expect the board question to shift by mid-year from whether people are using it to what it changed, which is the question this measurement exists to answer.
AI Value Measurement Review — we baseline the processes nobody has measured, design a rollout that produces a real comparison, and give your board three numbers instead of a percentage.
