Post-mortems on failed ERP programmes almost always blame requirements. The business did not know what it wanted. The design missed a critical process. Scope crept. Stakeholders disengaged. These findings are comfortable because they locate the failure at the beginning, where everyone involved has already moved on. The pattern in practice is different. Most large system programmes have a reasonable design. What they do not have is enough time to test it, because testing sits at the end of the plan and the end of the plan is where every earlier delay is absorbed. The arithmetic is predictable. Design runs four weeks long. Build absorbs a vendor delay. Data migration takes three attempts. The go-live date does not move, because the date was announced to the board, tied to a financial year, or committed in a contract. The only phase between the delays and the deadline is user acceptance testing, and it gets compressed from eight weeks to three. The defects that should have been found in those five weeks are found in production instead, by people trying to close the books.
Why the Cost Multiplies After Go-Live
A defect found in testing costs the time to reproduce it, fix it, and retest. A defect found in production costs considerably more, and the difference is not linear. In production, the defect has already produced bad data — wrong postings, incorrect tax calculations, mispriced orders, failed payments — which must be identified and corrected across however many transactions were affected. It has consumed support capacity while people work out whether it is a bug, a configuration error or a training issue. It has generated manual workarounds that persist long after the fix. It has damaged confidence in the system, which affects adoption of everything else. And the fix itself must be made under change control, in a live environment, with a regression risk that did not exist in a test system. The usual estimates put the multiplier somewhere between five and a hundred times depending on the phase and the nature of the defect. The precise figure matters less than the direction: compressing testing to save five weeks routinely costs several months of stabilisation.
The Six Ways Testing Fails
It is scheduled last and treated as a buffer. Every plan that places UAT immediately before a fixed go-live has already decided that testing will absorb the slippage. If the date cannot move, testing time cannot be protected unless it is protected explicitly. Business users are not released to do it. Testing is assigned to the people who best understand the process — who also have a day job that nobody backfilled. They test in gaps, superficially, under pressure to sign off. Un-backfilled testers is the most reliable single predictor of poor test coverage. Test data is not production-like. Clean, small, synthetic data passes. Real data — with historical oddities, incomplete records, unusual customers, legacy codes and the transaction from 2009 that nobody can explain — fails. A test cycle run on invented data validates the happy path and nothing else. Scenarios are written from the design, not from the business. The design document says how the process should work, so tests confirm that it does. The exceptions, the month-end variants, the intercompany cases, the credit note reversal, the partial delivery — the situations that make up a substantial share of real volume — were never in the design and therefore never in the tests. Integration and volume are tested inadequately or not at all. Individual modules pass. The end-to-end flow across systems, and the behaviour at realistic transaction volume at period end, is where production failures cluster and where testing is thinnest. And sign-off is a formality. Users are asked to approve after a compressed cycle where many scripts were not run and defects remain open. They sign because the alternative is to be the person who delayed the programme. The signature transfers accountability without transferring confidence.
What Good Testing Discipline Looks Like
Protect the test window contractually and in governance. Fix the testing duration, not only the go-live date. If earlier phases slip, the go-live moves. Making that rule explicit at the start is the only version that survives contact with pressure. Backfill the testers. People asked to test must be genuinely released from operational duties for the period, with cover arranged. Anything less produces attendance rather than testing. Use masked production data. Real volumes, real edge cases, real history, with personal and sensitive data masked. This single change surfaces more defects than any other testing improvement. Write scenarios from actual transaction history. Extract what really happened over a representative period — including the unusual ten percent — and build tests from that rather than from the process design. Test end-to-end and at volume. Full business cycles across integrated systems, plus a period-end simulation at realistic volume. Module-level testing does not predict either. Define entry and exit criteria numerically. Which scripts must pass, what defect severity is acceptable at exit, what coverage is required. Sign-off against criteria is meaningful; sign-off against a deadline is not. And rehearse the cutover more than once. Data migration, reconciliation, opening balances and the first close should be practised end to end, timed, with a tested rollback. Most go-live disasters are cutover failures, not functional ones.
Window
Protect testing from earlier slippage.
People
Arrange operational cover for testers.
Cases
Use masked production-like data and exceptions.
Flows
Test integrations, volumes, period-end and multiple entities.
Exit
Define coverage, pass and defect criteria.
Cutover
Rehearse migration, reconciliation, first close and rollback.
Qualitative summary of this article's source text, not a measured outcome or performance estimate.
Practical Guidance for ERP Test Strategy
- Fix the testing duration before fixing the go-live date. Whichever is protected first is the one that survives.
- Backfill testers with real operational cover. Testing done in the margins of a day job is not testing.
- Test with masked production data at realistic volume. Synthetic data validates the happy path and hides everything that matters.
- Derive scenarios from transaction history, including exceptions. The design document describes the intent, not the traffic.
- Make integration and period-end the priority test areas. This is where production failures concentrate.
- Set numeric entry and exit criteria for each cycle. Coverage percentage, script pass rate and permitted open-defect severity.
- Rehearse cutover and the first close, twice, with a rollback plan. Time it; the dress rehearsal is where the real schedule emerges.
- Escalate compression as a risk decision, not a scheduling detail. Someone accountable should knowingly accept a shorter test window, in writing.
The Regional Angle
Testing requirements for Gulf implementations include several things that standard global test plans routinely omit, and each of them tends to be discovered in production. Statutory outputs must be tested against the actual regulator, not the specification. ZATCA e-invoicing in Saudi Arabia involves structured invoice formats, cryptographic stamping and clearance or reporting flows; testing must include real submission through the applicable sandbox and rejection handling. WPS payroll file generation must be validated against the bank's accepted format for each entity. VAT and corporate tax returns must be produced and reconciled. A configuration that looks correct and is rejected by the regulator's system is a production incident by definition. Arabic and bilingual output needs explicit test cases. Right-to-left rendering, Arabic invoice layout, dual-language statements, Arabic characters in names and addresses flowing correctly through to printed and electronic documents, and sort order behaviour in bilingual master data. Systems that pass every English test case routinely fail these, and the failures appear on customer-facing documents. Gratuity and leave calculations need scenario coverage. End-of-service benefit calculation varies by jurisdiction, length of service, contract type and reason for separation. These rules have genuine complexity and real financial consequence, and they are usually tested with two simple cases. Multi-entity and intercompany flows must be tested as a group, not per entity. Cross-entity sales, shared service recharges, currency translation and elimination behave differently in combination. Per-entity testing passes while the consolidation fails. Calendar assumptions need validation. Differing weekend days across countries, Ramadan working hours, and public holidays that shift with the lunar calendar affect period definitions, service level calculations, payroll cut-offs and scheduled jobs. And tester availability is genuinely harder here. Small local finance teams, staff on leave outside the country for extended periods, and visa-linked turnover mid-programme all reduce the pool of people who can test. This must be planned for rather than discovered.
What Changed, and What Did Not
The technical side of testing improved considerably after 2014. Automated regression suites made repeat testing cheap, which matters enormously for cloud ERP products that update on a vendor schedule rather than a customer one. Continuous delivery practices from software engineering reached enterprise systems, shifting the model from a single large test event to ongoing validation. Cloud environments made it practical to spin up realistic test instances without a hardware business case. What did not change is the organisational cause. Testing is still scheduled last, still staffed by people who were not released, still compressed when earlier phases slip, and still signed off under deadline pressure. That is a governance failure, and no tooling fixes it. AI is starting to help in specific places: generating test scenarios from transaction history and process documentation, producing test data with realistic distributions, identifying which regression tests are relevant to a given change, and triaging defect reports. Early evidence is that this raises coverage meaningfully for the same effort. It also introduces a new testing obligation that most programmes have not absorbed. When AI features are embedded in the ERP — automated coding suggestions, anomaly flags, document extraction, natural-language reporting — they cannot be tested as deterministic functions. They need accuracy thresholds, sampled human review, defined behaviour when confidence is low, and monitoring in production rather than a one-time pass. Very few test strategies have been updated for that, which means the next generation of post-go-live surprises is already being built in.
Common Questions
Why do ERP projects fail at testing rather than design?
Because testing is scheduled immediately before a fixed go-live date, so every earlier delay is absorbed by compressing the test window. The design is usually adequate; the validation of it is what gets cut.
How much more expensive is a production defect?
Estimates commonly range from roughly five to a hundred times the cost of finding the same defect in testing, depending on phase and severity. The difference comes from corrupted data requiring correction, support load, manual workarounds, lost confidence and the risk of fixing under change control in a live system.
What is the single highest-value testing improvement?
Using masked production data at realistic volume, with scenarios derived from actual transaction history rather than the design document. It surfaces the exception cases that cause most production failures.
What do global test plans miss in the GCC?
Regulator-facing submission testing such as ZATCA e-invoicing and WPS payroll files, Arabic and bilingual document rendering, gratuity and leave calculation scenarios, group-level intercompany and consolidation flows, and calendar logic for differing weekends, Ramadan hours and lunar holidays.
Test Strategy Review — Outpace protects the test window, builds scenarios from your real transactions, and rehearses the close before it counts.
