In 2017 the phrase "cognitive automation" was everywhere in back office procurement. Vendors promised software that would read documents, understand context, make judgements and learn from corrections — replacing not just the keystrokes of clerical work but the reasoning behind it. Analyst decks showed a straight line from rule-based robotic process automation to genuine machine judgement, with the transition already underway. What was actually deployed and running in production, at most organisations, was rule-based automation with a pattern-matching layer on top. It read invoices well, provided the invoices looked like the ones it had seen. It routed emails accurately, until the phrasing changed. It handled the common cases and escalated everything else, which turned out to be a larger share of volume than the business case assumed. That gap between the pitch and the deployment is worth studying carefully, because the same gap reappeared with every subsequent wave — and the organisations that learned to measure it are the ones now getting real value from genuinely capable models.
Where the hype cycle went wrong
Three specific errors recurred. Confusing perception with judgement. Extracting fields from an invoice is a perception task and was largely solvable. Deciding whether a disputed invoice should be paid requires knowing the contract, the delivery history, the commercial relationship and what the account manager promised on a call. Vendors demonstrated the first and priced the second. Underestimating the exception tail. Back office processes follow a long-tail distribution: a modest set of patterns covers most volume, and the remainder is genuinely varied. Automating the common cases is valuable, but it does not scale linearly — and since exceptions consume disproportionate handling time, removing the easy 70% of volume might only remove 40% of the effort. Business cases built on volume percentages consistently overstated savings for exactly this reason. Ignoring the training data problem. A learning system needs labelled examples of correct decisions. Most back office processes had no such record — outcomes were captured, reasoning was not. Organisations discovered they could not train the intelligent system because they had never written down why anyone did anything.
What was genuinely working
Stripped of the marketing, several things did deploy successfully and are worth remembering because they set the pattern for what works. Document extraction with confidence scoring, where high-confidence results processed automatically and low-confidence ones went to a human queue. Classification and routing of inbound requests, which removed a real triage burden. Validation and matching against master data, catching errors earlier than a downstream reconciliation. Prediction of processing outcomes — which invoices will fail, which requests will escalate — used to prioritise rather than to decide. The common feature is that all four keep a human in the decision and use the machine for speed and coverage. The deployments that failed were the ones that tried to remove the human entirely from a process where the consequences of being wrong were material and the exception rate was unknown. There is a design lesson in that which has outlasted the terminology: build the confidence threshold and the human queue first, then move the threshold as evidence accumulates. Organisations that started at full automation and retreated lost the trust of the operational team, and that trust is harder to rebuild than the system.
Practical Guidance for Automation Reality Check
- Ask for a production reference at your volume and exception profile, not a demo. Demos run on curated documents; ask what the straight-through rate was in month six.
- Measure effort by exception, not volume by case. Removing 70% of transactions may remove far less than 70% of the work, and the business case depends on which you counted.
- Separate perception tasks from judgement tasks in your scope. Reading and classifying are solved problems; deciding is not, and pricing frequently blurs the two.
- Check whether you have training data before buying a learning system. Outcomes without recorded reasoning cannot train a model to reason.
- Design the human queue and confidence threshold on day one. Start conservative and loosen with evidence; the reverse sequence destroys operational trust.
- Instrument straight-through rate, exception reasons and rework from the start. Without those three measures you cannot tell whether the system is improving or the volume mix changed.
- Fix the upstream input before automating the downstream handling. Structured submission at source beats clever extraction from unstructured documents every time.
- Cost the maintenance honestly. Automation degrades as formats, rules and counterparties change; someone owns that forever.
The Regional Angle
Gulf back office operations present a harder version of this problem than the vendor case studies assume, and a few of the differences are decisive. Bilingual document handling is the first. Invoices, purchase orders, delivery notes, contracts and government correspondence arrive in Arabic, English or both, frequently in the same document, with mixed right-to-left and left-to-right layout. Extraction accuracy on Arabic text — particularly handwritten annotations, stamps, and scanned faxes of poor quality — has historically lagged well behind English, and vendors quoting accuracy figures from Western deployments are quoting numbers you will not reproduce. Insist on a pilot using your own documents, with Arabic-language samples included, before signing anything. The manual document culture is the second. Physical stamps and wet signatures still carry authority in many regional transactions and government interactions. Paper-based approval chains, hand-annotated invoices, and PDFs that are photographs rather than digital documents remain common. That raises the perception difficulty while simultaneously increasing the potential value, because the manual handling cost being displaced is genuinely high. Third, the regulatory layer has changed the economics in a useful direction. Saudi e-invoicing clearance and the broader move to structured electronic invoicing across the region converts a share of inbound documents into machine-readable formats by regulation, which removes the extraction problem entirely for those transactions. The strategic implication is important: rather than investing heavily in extracting data from unstructured regional documents, the better sequencing for many organisations is to push counterparties toward structured submission — supplier portals, e-invoicing channels, standardised templates — and automate what arrives cleanly. Finally, the labour arithmetic differs. Regional shared service centres in Dubai, Riyadh and Cairo were built partly on favourable processing costs, which changes the payback calculation for expensive automation platforms relative to higher-wage markets. But the other side of the regional labour picture cuts the other way: employment-linked residency makes turnover costly and disruptive, and tacit process knowledge walks out with each departure. The strongest regional case for automation is usually continuity and auditability rather than headcount reduction — and framing it that way tends to get a better hearing internally, while also being more honest about what these systems reliably deliver. Nationalisation targets under Emiratisation and Saudisation add a further consideration: automation that removes entry-level clerical roles may conflict with commitments to create local employment, which is a genuine strategic tension rather than a compliance footnote.
The objection worth taking seriously
The strongest challenge to this sceptical framing is that the sceptics were eventually wrong about the trajectory, even though they were right about 2017. Large language models delivered much of what cognitive automation promised, and they did it within a few years of these deployments disappointing. Systems can now read a contract and answer questions about it, handle documents in formats they have never seen, work across languages including Arabic at genuinely useful accuracy, and produce reasoned drafts of decisions with explanations. An organisation that concluded in 2017 that this category was overhyped and stopped paying attention made a defensible judgement about the technology of that moment and a poor one about the direction. Timing scepticism is not the same as capability scepticism, and it is worth being explicit about which one you hold. There is also a fair criticism of the human-in-the-loop design principle. Keeping a human in every decision caps the benefit at review speed, and review is not free — reviewers become rubber stamps when volumes are high and the system is usually right, which produces the worst combination: the cost of human oversight without its protection. Automation bias is well documented, and a review step that nobody genuinely performs is a governance illusion. The better designs sample rather than review everything, monitor the sampled error rate, and escalate on drift, which is a meaningfully different control from "a human approves each one". The defensible position: the 2017 disappointment was a lesson about evaluating claims against deployed evidence, not a lesson that automation of judgement is impossible. Apply the same discipline now — production references, exception measurement, honest maintenance costs — to technology that has genuinely improved.
Common Questions
Why did cognitive automation underdeliver in 2017?
Because the deployed technology handled perception well and judgement poorly, exception tails were larger than the business cases assumed, and most organisations lacked the labelled decision data that learning systems required.
What is the right metric for automation success?
Straight-through rate with a stable exception profile, plus total handling effort before and after. Volume automated is a vanity metric, because the automated volume is usually the cheap volume.
Should we automate a process or fix it first?
Fix, simplify and where possible restructure the input first. Automating a bad process makes it faster and harder to change, and structured submission at source removes more cost than any extraction technology.
Has AI closed the gap now?
Largely, on the perception and drafting side, and partially on judgement. Current models read varied document formats, handle Arabic and mixed-script content far better, extract from layouts they have not seen, and can produce a reasoned recommendation with its justification — which addresses the 2017 failures directly. Three things have not changed. The exception tail still exists, and models fail differently: rather than escalating what they do not understand, they produce a plausible answer, which means confidence calibration and sampling-based quality monitoring matter more than before, not less. Auditability is now the binding constraint in regulated back office work — you need to show why a payment was approved, and a model's explanation is not the same as an audit trail of the rule that applied. And accountability is unchanged: someone owns the outcome regardless of what produced it. The practical pattern that works is a model doing the reading and drafting, deterministic rules enforcing the controls that matter, and sampled human review with monitored error rates.
Automation Reality Check — ask for the month-six straight-through rate at your exception profile, and measure effort removed rather than volume automated.
