Every enterprise software roadmap presented this month contains the same three letters. Retrieval augmented generation is the answer vendors give when asked how their assistant will know anything about your business, and for document-heavy work the answer is a good one. Applied to an ERP, it needs translating. A RAG-powered ERP is not one thing, and the version being demonstrated is usually not the version that would answer the questions your controller actually asks.
Retrieval augmented generation was built to find a paragraph. An ERP does not contain paragraphs. It contains three hundred tables and a join
Document retrieval works because prose describes itself. Split a policy into chunks, embed each chunk, and a question about expense limits lands near the paragraph about expense limits, because the words are in it. Now embed a ledger line: 4,318.00, AED, 30-Nov, vendor 10042, cost centre 220, account 5310. There is no semantic content to match against. Nearest-neighbour search over transactional rows returns nothing useful, and the more rows you index the worse it gets. The information in an ERP lives in relationships and definitions, not in the text of the records.
Three things are being called RAG, and only one of them is
Retrieval over the unstructured periphery. Contracts, purchase order attachments, vendor correspondence, policies, approval memos, the notes field nobody cleans. This is genuine document RAG, it works today, and it is where most of the near-term value sits — precisely because this material has never been searchable. Retrieval over schema and definitions, then generated query. The user asks a question, the system retrieves the relevant field descriptions, code lists and metric definitions, writes a query, and runs it. The retrieval is over documentation; the answer comes from the database. This is the architecture most "ask your ERP" demos are running, and calling it RAG obscures where it fails. The hybrid, which is what actually survives production. Retrieve the definition and the governing policy, execute a deterministic query for the numbers, then have the model narrate the result and cite what it used. The model never invents a figure. It explains one. If a vendor cannot tell you which of these three you are buying, you are watching a demo rather than evaluating a system.
Retrieval quality is capped by how well your estate is described
This is the part that gets skipped, and it is the whole project. The corpus worth indexing is not your transaction history. It is the layer of description around it: what each custom field means and who owns it, the code lists and what the codes stand for, approval thresholds and delegation rules, the close calendar, the narrative behind the chart of accounts, and — most valuable and least available — the record of why configuration decisions were made. Most organisations discover at this point that this material does not exist in writing. That discovery is not a reason to stop. It is the project.
Entitlements have to be applied at query time
ERP data is segmented by company, cost centre, plant and role in ways that document repositories rarely are. The tempting shortcut is to build a separate index with its own access rules, which immediately becomes a second entitlement model that drifts from the first. The workable pattern is to resolve permissions when the question is asked, against the source system's own authorisation, and to filter retrieved context accordingly. It is slower. It is the only version that stays correct after the next reorganisation.
Evaluate by question type, and watch the refusals
Sort real questions into four kinds: lookups, aggregates, explanations and exceptions. Explanations are where retrieval shines — why is this invoice blocked, what does this status mean, which policy governs this approval. Aggregates must be deterministic, always, with no generative arithmetic anywhere in the path. Lookups are easy. Exceptions are hard, because answering them requires process context that lives in people. Then measure the refusal rate rather than only accuracy. A system that declines to answer when it lacks grounding is usable by a finance team. A system that produces a plausible number every time is not, and you will find that out during an audit.
| Question type | Grounding and evaluation |
|---|---|
| Lookup | Authorised source record and explicit entity scope. |
| Aggregate | Governed query and tested metric definition, not generated arithmetic. |
| Explanation | Source policy/definition and a cited, bounded response. |
| Exception | Process context, ambiguity handling and a refusal or human route. |
Qualitative summary of this article's source text, not a measured outcome or performance estimate.
Practical Guidance for RAG ERP Implementation Consultation
- Start with the unstructured periphery, where retrieval genuinely works.
- Never let a model compute a figure; retrieve definitions, execute queries.
- Build the description layer first — fields, codes, policies, decisions.
- Resolve permissions at query time, against the source system.
- Classify questions into four types and evaluate each separately.
- Require citation of the source record or document on every answer.
- Track refusal rate alongside accuracy, and reward refusals.
- Budget re-indexing as an operating cost, not a one-off.
The Regional Angle
The first obstacle for most groups here is that the corpus you would index does not exist, and the people who could write it do not work for you. Regional ERP estates are overwhelmingly partner-implemented, often across several partners and several waves, and the documentation deliverable was frequently the first thing cut when the go-live date moved. What the retrieval layer needs — why this field was added, what this status code means in practice, which of the four approval routes is actually used — exists in the working memory of two or three consultants, some of whom have moved on. Before scoping any retrieval project, read your support contract and establish whether the partner owes you documentation and in what form. Where they do not, the fastest route is structured interviews, recorded and transcribed, turned into indexed descriptions. It feels like an odd way to start an AI project. It is the highest-return week in the whole programme, and the asset survives every model you subsequently replace. The second is that the highest-value content to index is also the most dangerous to get wrong. Questions about VAT treatment, zero-rating, free-zone versus mainland status, e-invoicing requirements and wage protection obligations are exactly what staff will ask an assistant, because the rules are recent, they change, and the answers are hard to find. An assistant that states a tax treatment confidently and incorrectly does not create an inconvenience; it creates a filing position someone relied on. Index the official published text rather than a summary or a partner's slide, attach a version and an effective date to every rule document, and require the assistant to cite the rule and its date whenever it answers a statutory question. If you cannot meet that standard, exclude tax and regulatory topics from generative answering entirely and route them to a human. That is a legitimate design decision and a defensible one. The third is scoping. In a ten-entity group the same customer, the same vendor, the same cost centre name and the same project code appear repeatedly across legal entities with different charts of accounts and, increasingly, different functional currencies. "What did we spend with this supplier this year" has ten different correct answers and one wrong one, which is the sum. Retrieval systems that treat entity as an ordinary filter will return confidently wrong results, because the wrong entity's data is semantically indistinguishable from the right one's. Make entity scope an explicit, required parameter on every query path, have the assistant state which entity it answered for in the response, and test the ambiguous cases deliberately during evaluation. Groups that skip this do not discover the problem in testing. They discover it when two managing directors receive different numbers for the same question.
The objection worth taking seriously
The sharpest objection is that all of this is a wrapper with a twelve-month shelf life. The major ERP vendors have already announced assistants; the platform vendors will ship retrieval as a feature of the database and the collaboration suite; and anything you build now with an embedding pipeline and a vector store will be superseded by something native, supported and cheaper. Building a retrieval stack in 2024 looks a great deal like building a portal in 2004. The prediction is probably right, and it should change what you build. The plumbing — the embedding pipeline, the vector store, the orchestration — is the part that will be commoditised, and you should not fall in love with it. Buy it, rent it, keep it thin, and assume you will throw it away. What will not be commoditised is the description layer. No vendor can write down what your custom status codes mean, which of your approval routes is real, how your entities relate, or why the previous controller created three separate accounts for what looks like one thing. Every assistant you ever deploy — native, third-party or whatever arrives in three years — will need that material, and the organisations that spend this year producing it will be able to switch tools in a fortnight. The ones that spend the year evaluating vector databases will still have an undocumented ERP. Build the asset. Rent the model.
Common Questions
Can we point retrieval straight at the production database?
For documents, yes, with permission filtering. For transactional tables, no — you want generated queries against a governed layer, not similarity search over rows.
How often does the index need rebuilding?
Documents on change. Descriptions and code lists whenever configuration changes, which means tying re-indexing to your change process rather than to a schedule.
Does this replace reporting?
No. It answers questions about reports, definitions and exceptions. The numbers should still come from the same governed queries the reports use.
What should we expect over the next twelve months?
Expect every major vendor to ship a retrieval-based assistant inside the suite during the year, with the initial releases confined to help content and the unstructured periphery rather than the transactional core. Expect the first serious incidents to involve entitlements — an assistant surfacing something to someone who should not have seen it — and expect that to slow enterprise rollouts more than accuracy will. Expect pricing to move from per-seat toward consumption as re-indexing and context costs become visible. And expect the competitive question by year end to be about who has documented their estate, not who has chosen the better model.
RAG ERP Implementation Consultation — we build the description layer your assistant will need, whichever vendor eventually supplies the assistant.
