Editorial context. The source date is retained. The opening dashboard scenario is illustrative, not a documented client case. The later discussion of AI reflects retrospective commentary, not technology adoption in April 2007.
The dashboard is green. Ninety-eight percent of invoices processed within the agreed window. Query response inside four hours. Availability above target. Every service level met for eleven consecutive months. The finance director, meanwhile, thinks the service is poor, the business units have quietly rebuilt their own teams, and nobody can explain why suppliers still call about unpaid invoices. This is SLA theater, and by 2007 it had become the dominant failure mode of business process outsourcing. India's ITES-BPO export sector was heading past $8 billion in revenue that year, first-generation contracts signed around the turn of the decade were coming up for renewal, and buyers were discovering an uncomfortable pattern: contracts were being met and expectations were not.
Where the Bad Metrics Came From
Business process SLAs were inherited, almost unchanged, from IT infrastructure outsourcing. That lineage explains nearly every defect in them. Infrastructure contracts measured availability, response time and resolution time, because for a data centre those things genuinely describe the service. Applied to an invoice, a payroll run or a customer query, the same structure measures the vendor's internal activity and says nothing about whether the business outcome occurred. Four design flaws followed directly. The measurement boundary excludes the waiting. The clock starts when the item enters the vendor's queue and stops when it leaves. Time spent waiting for an approval, a missing purchase order or a clarification from the business is excluded. A process that takes six weeks end to end can be reported as two days of service delivery, and both numbers are honest. Averages hide the failures that matter. A ninety-eight percent completion rate against a monthly volume of forty thousand transactions leaves eight hundred failures. Those eight hundred generate the complaints, the escalations and the executive perception of the service. The two percent is the service, as far as your internal customers are concerned. Activity is measured instead of outcome. Calls answered rather than problems solved. Invoices processed rather than suppliers paid correctly. Tickets closed rather than issues resolved. Each activity metric can be improved by behaviour that makes the outcome worse. The vendor measures itself. Data comes from the vendor's workflow tool, on the vendor's definitions, reported monthly by the vendor. Not dishonestly, in most cases — but every ambiguity in the definition resolves in the reporter's favour, consistently, for years.
Why Service Credits Do Not Fix It
The standard remedy is a credit: miss the target, forfeit a percentage of the monthly fee. It sounds like accountability. Structurally, it is a price list for underperformance. Credits are typically capped at a small share of monthly charges, which is trivially small relative to the cost of a late payment run, a mis-stated ledger or a lost customer. Once a vendor calculates that the credit costs less than the resourcing required to hit the target, the credit becomes an operating expense and the target becomes optional. Credits also create a specific pathology: with per-metric credits, a vendor under pressure will protect the measured metrics by moving resources away from everything unmeasured. Service quality in the gaps between KPIs degrades exactly as fast as the measured metrics improve. And there is the volume problem. Most contracts price per transaction, which rewards the vendor for volume. Eliminating unnecessary transactions — the single most valuable thing that could happen to your back office — reduces the vendor's revenue. Nobody should be surprised when process improvement proposals stop arriving.
What to Measure Instead
A working performance framework has a small number of metrics that describe outcomes the business recognises.
- End-to-end cycle time, measured from your systems, from the moment the business initiates the request to the moment the outcome is complete. Include every queue, whoever owns it. Compare this with the vendor-reported figure. The difference depends on which queues and waiting periods each measure includes; do not assume a standard multiplier.
- First-time-right rate. The proportion of transactions completed without rework, correction or a follow-up query. This is the closest single proxy for real quality.
- Tail performance, not averages. Report percentiles using the same population and clock. A target of "ninety-five percent within two days" imposes a tighter time threshold than "ninety-fifth percentile under four days." Neither target limits how long the slowest five percent may take. Track those exceptions separately.
- Exception volume and root cause. Exceptions are where cost and dissatisfaction live. Track their number, their causes and who owns each cause. Frequently the answer is the retained organization, which is the point.
- Cost per business outcome, not cost per activity. Cost per supplier paid accurately, not cost per invoice touched.
- Internal customer experience, gathered from the business units directly, quarterly, and reported alongside the operational metrics without the vendor filtering it. Six metrics with consequences beat sixty in a monthly pack that nobody reads.
| Activity-only view | Outcome or context to retain |
|---|---|
| Receipt or queue clock | Start and stop points, including waits experienced by the requester. |
| Average response | The slow tail and the work outside the average. |
| Items processed | Accuracy, resolution and supplier or customer outcome. |
| Provider-produced score | A retained customer evidence trail and joint review. |
Qualitative summary of this article's source text, not a measured outcome or performance estimate.
Governance That Works
- Measure from your systems, not the vendor's, wherever technically possible. Where it is not, agree the extraction method in the contract and audit it.
- Define the clock precisely — including what stops it, and who is accountable for each pause. Most disputes are definitional, not performance-related.
- Replace credits with earn-back and gainshare. Let the vendor recover credits through sustained improvement, and share the savings from volume they help eliminate. Align the commercial model with the outcome you actually want.
- Make the quarterly review substantive. Root causes, improvement commitments with owners and dates, and progress against last quarter's commitments. A review that reads a dashboard aloud is the purest form of the theater.
- Include a sustained-failure exit right, distinct from the credit regime. A vendor that can only lose money has less discipline than one that can lose the contract.
- Review the SLA when the process changes. Volumes shift, scope creeps and targets set three years ago stop describing the work. Stale targets are met automatically, which is why nobody proposes revising them.
The Part Contracts Cannot Fix
An SLA cannot repair an unclear scope boundary, a process that was broken before it was transferred, or a retained organization that never decided who owns exceptions. Those are design problems, and they show up as performance problems only because the performance report is the one place anybody looks. The question is now getting sharper rather than softer. As AI handles more transaction processing, throughput and turnaround targets are becoming trivially easy to meet, which will make green dashboards even less informative. The metrics that survive are the ones measuring accuracy, exception handling and end-to-end outcomes — the things that were always worth measuring and were always the hardest to put in a contract.
Questions Operators Ask
What makes an SLA meaningful rather than theatrical?
Outcome-based metrics measured from the buyer's systems, end-to-end clocks that include every queue, tail percentiles instead of averages, and consequences that exceed the cost of underperforming.
Why are all our SLAs green when the business is unhappy?
Almost always because the measurement boundary excludes waiting time owned by other parties, and because averages conceal the minority of transactions that generate all the complaints.
Are service credits worth negotiating hard?
Only to a point. Credits are capped and rarely approach the business cost of failure. Negotiating measurement definitions, end-to-end scope and exit rights delivers far more leverage than negotiating credit percentages.
How many KPIs should a BPO contract have?
Enough to describe the outcomes the business cares about and few enough that an executive can review all of them seriously in one meeting. In most contracts that is between five and eight.
SLA Redesign Workshop — Outpace rebuilds outsourcing performance frameworks around end-to-end outcomes measured from your systems, replaces credit regimes with commercial terms that reward improvement, and gives your quarterly reviews an agenda worth attending.
