ERP / Source date:

Master Data Governance Is the Prerequisite for Enterprise AI

Models inherit every inconsistency in customer, item, and vendor records and amplify them at scale.

Illustration of a data owner comparing duplicate supplier folders with an authoritative master-record binder.

Every artificial intelligence proposal circulating in enterprise IT this quarter carries the same unpriced dependency. The business case assumes the model will answer questions about customers, suppliers, products and costs. It does not mention that the records describing those things are, in most organisations, partly duplicated, partly abandoned and owned by nobody in particular. Master data governance has been a slow-moving programme for two decades, funded reluctantly and rarely finished. The arrival of generative tools has changed its status from hygiene to dependency, and the change happened faster than most data teams were prepared for.

Bad master data used to produce a wrong report, and someone usually caught it. It now produces a fluent, confident paragraph, and nobody does

That is the whole argument, and it is worth stating plainly before the frameworks arrive.

Three reasons artificial intelligence changes the calculus

Consumption scales. A monthly report is read by a handful of people who have worked with that data for years and can tell when something looks wrong. An assistant answers hundreds of unrehearsed questions for people who have no such instinct and no reason to doubt the answer. The sanity check disappears. The analyst who quietly knew that a supplier had been dormant since 2019, that one plant codes things differently, and that the September figure is always restated — that person was a control. They are not in the loop any more. Duplicates become arithmetic. Three records for one customer were a nuisance in a list and are a wrong number in an answer. Aggregation is where fragmented master data stops being untidy and starts being false.

The objects that matter, and the two nobody lists

Customer, supplier, item and employee are the obvious four. Fix them in that order of business impact, not alphabetically. The two that get omitted are the ones that break answers most often. Organisational hierarchy — cost centres, profit centres, reporting lines — because every rolled-up figure depends on it and it is usually maintained in a spreadsheet by one person. And the mapping layer between local charts of accounts and the group structure, because that is where an answer becomes wrong without becoming obviously wrong. Hierarchies deserve particular attention. A duplicate customer produces a visibly odd result. A misassigned cost centre produces a perfectly plausible one.

Governance is four things, and none of them is a product

An accountable owner for each object, named, with time allocated. A written definition of what constitutes a duplicate, because that is a business judgement and not a technical one. A creation and change process with validation enforced at the point of entry. And a published measurement, so the state of the data is visible without anyone having to ask. Organisations buy a master data management platform and skip all four, then conclude that master data management does not work.

Four numbers, published monthly

Duplicate rate per object. Completeness against required fields. Staleness — records untouched for eighteen months that remain active. And orphan rate, meaning live transactions pointing at records that should have been retired. Put them on one page, send them to the same distribution every month, and name the owner next to each. Most of the improvement comes from the publishing, not from the tooling.

Clean less than you think

The instinct is a full cleanse. It is almost always the wrong shape of project, because it spends the budget on records that no transaction has touched in years. Run the transaction history instead. Identify the master records actually used in the last eighteen months — typically a quarter to a third of the file — and clean those properly. Deactivate the rest rather than cleansing them, keeping them readable for audit. You will have finished the part that matters before a comprehensive programme would have completed its scoping phase. Then move the effort to prevention: one creation path rather than five, a duplicate check before save, mandatory fields actually mandatory, and separation between who creates a record and who approves it.

Resolve the ambiguity your use case touchesQualitative scope from the article. No fixed cleanup period or improvement percentage is promised.
  1. Scope the use case

    Identify the records and reporting hierarchies its answers depend on.

  2. Assign business ownership

    Define duplicates and settle which records represent the same real entity.

  3. Validate changes

    Use an approved creation path and controls at entry.

  4. Track quality

    Publish duplicate, completeness, staleness and orphan measures with named owners.

Qualitative summary of this article's source text, not a measured outcome or performance estimate.

Practical Guidance for MDM Readiness Assessment

  • Name one accountable owner per object, with allocated time.
  • Define what counts as a duplicate in writing, as a business rule.
  • Clean only the records transactions touched in eighteen months.
  • Deactivate the remainder rather than cleansing it.
  • Enforce validation at creation, not in a monthly cleanup job.
  • Separate record creation from record approval, with an SLA.
  • Publish four quality numbers monthly with owners named.
  • Fix hierarchies before attributes — rollups fail silently.

The Regional Angle

The first regional complication is that your counterparties are frequently yourselves. Groups here trade internally at volume: the same company is a customer of one entity, a supplier to another and a related party to a third, and the same individual may sit on both sides of a transaction. Since corporate tax arrived, that is no longer just a consolidation inconvenience — related-party transactions carry documentation and arm's-length pricing obligations, and the disclosures are prepared from the same master data that was never designed to carry the flag. Add related-party status as a governed attribute on the customer and supplier objects, with a maintenance owner in tax rather than in finance operations, and populate it at record creation rather than reconstructing it at year end from memory and a shareholding chart. Groups that leave this to the annual scramble discover that the answer depends on who is asked, which is precisely the position transfer pricing documentation is meant to prevent. The second is the item master in trading, contracting and distribution businesses, where a single physical product carries four identities: the supplier's part number, the customs classification code, the line reference in a project bill of quantities, and the descriptive name a storekeeper actually uses. Historically only the last one was maintained with any care. That has become untenable, because the customs code now drives duty treatment, the tax treatment flows into e-invoicing fields that are validated by the authority rather than by you, and the unit of measure has to match across purchase, stock and sale or the reconciliation fails. Make classification code and unit of measure governed, mandatory fields with a named owner in logistics or trade compliance, and audit them against a sample of actual declarations. An item master that was merely untidy is now a filing risk. The third is employee identity, which in most regional organisations is spread across four systems that disagree. Human resources holds one record, payroll holds another, the government relations function maintains the visa and permit file, and the access control system holds a third version created by whoever set up the laptop. The authoritative document is frequently a scanned passport in a shared folder. Decide which system is the register of record for employee identity, key it to the national identity number rather than to a name that will be spelled three ways, and treat residency permit, labour card and passport expiry dates as governed master data attributes that trigger workflow rather than as scanned attachments. The payoff is not only data quality. It is that an assistant asked a question about headcount, cost or entitlement gives the same answer to everyone who asks it.

The objection worth taking seriously

The strongest objection is that this is the oldest stalling tactic in enterprise IT. Data quality programmes have been used to defer value delivery for twenty years, they are famously hard to finish, and the business never sees an outcome it can point to. Worse, the argument may be obsolete: modern models are genuinely good at reconciling messy input. They cope with inconsistent formatting, spelling variation and partial records far better than the rule-based matching that preceded them. Insisting on clean master data before touching artificial intelligence risks spending two years on hygiene while competitors ship something useful. That critique is largely fair, and anyone proposing a comprehensive cleanse before a first use case should be refused. Perfect data is not a prerequisite and has never been achievable. The distinction that matters is between messy and ambiguous. Messy is a formatting problem, and models handle it well — inconsistent addresses, mixed date formats, a name written four ways. Ambiguous is a question of fact about the business: which of these three records is the real customer, whether these two entities are the same counterparty, which cost centre this sits under. No model can resolve that, because the answer is not in the text; it is a decision somebody in the organisation has to make and record. Scope the governance narrowly to the objects and hierarchies your first two use cases actually touch, resolve the ambiguity there, and leave the rest messy. That is a quarter of work, not a programme, and it is the difference between an assistant that is useful and one that is quietly wrong.

Common Questions

Do we need a dedicated master data platform?

Usually not at first. Ownership, a definition, validation at entry and a monthly measurement will carry most organisations a long way. Buy tooling when the manual process is demonstrably the constraint.

Who should own master data?

The business function that suffers most when it is wrong, supported by IT. Supplier data belongs with procurement, not with the integration team.

How do we justify the spend without an AI project attached?

Use the numbers you already lose: duplicate payments, credit exposure to what turns out to be one counterparty, rework at close. Those predate artificial intelligence and are easier to quantify.

What should we expect over the next twelve months?

Expect the first wave of assistant deployments to stall on exactly this, and expect the failure to be reported as a model problem rather than a data problem. Expect vendors to market matching and enrichment capabilities heavily during the year — useful for the messy half, irrelevant to the ambiguous half. Expect regulatory pressure, particularly from e-invoicing and transfer pricing documentation, to do more for master data budgets than any internal business case has managed in a decade. And expect the organisations that quietly fixed their hierarchies this year to be the ones whose second-wave projects work.


MDM Readiness Assessment — we scope the governance to the records your first use cases actually touch, and leave the rest alone on purpose.

Continue reading

Talk to OPS

Start with the operating problem.