ERP / Source date:

ChatGPT for ERP: Natural Language Queries Replace Reports

Natural language ERP interfaces in 2023 finally delivered on a decade of promises — enabling any employee to query business data in plain English without SQL, report builders, or analyst intermediaries.

Illustration of reviewing a revenue metric's definition, entity scope, currency, exclusions and owner before agreement.

Every ERP vendor is now demonstrating the same thing. Someone types a question in plain English — which customers are over their credit limit, what did we spend on freight last quarter, show me open purchase orders past their delivery date — and a number appears without anyone opening a report. The demonstration always works. It works because the demonstration database has a few dozen tables, one company code, one currency and exactly one definition of revenue. Yours does not.

A model that writes perfect SQL against a schema nobody agrees on will produce a perfectly formed wrong answer, quickly and confidently

Translating a question into a query is the part that is close to solved. Current models do it well against clean schemas, and they will do it better in six months. The unsolved part is everything the query needs to know that is not written down anywhere. Which of the four tables holding sales figures is the one finance uses. Whether revenue is gross or net of credit notes. Whether the discontinued company code from the 2019 acquisition should be included. Whether cancelled lines count. Whether the intercompany entries are eliminated at source or in consolidation. That knowledge exists in your organisation. It lives in the heads of three people in finance and in the formulas of a spreadsheet one of them maintains. It is not in the database, so it is not available to the model.

Three failure modes, in order of how much damage they do

The definition problem. Revenue, margin, headcount, open orders and days sales outstanding each have several defensible definitions inside a single company. Ask a person and they will ask you which one you mean. Ask a language model and it will choose, silently, and never mention that a choice was made. The silent filter. Real queries in a real ERP require exclusions that nobody states out loud: test entities, the legacy company code, intercompany traffic, reversed documents, the plant that closed. A generated query includes everything unless told otherwise, and the resulting number is wrong in a direction nobody can spot. Plausibility. This is what makes the category different from a broken report. A failed report throws an error. A wrong generated answer returns a number of the right magnitude, correctly formatted, with a confident sentence explaining it. There is nothing to notice.

Permissions are the reason these features stay in preview

ERP authorisation is not a table-level matter. It is object-level, field-level and organisational: this user sees this company code, that cost centre, those plants, and not the payroll fields at all. A conversational layer that connects with elevated rights, runs the query and then filters the result is a data leak with extra steps, because aggregates leak too. A regional manager who can ask what the total payroll cost is for a department of four people has learned four salaries without seeing a single restricted field. The requirement is that the query executes with the asking user's own authorisations, inside the ERP's permission model rather than beside it. That is difficult, it is the main reason these capabilities ship as restricted previews, and it is the first question to ask any vendor.

The semantic layer is the actual project

The useful work is not the interface. It is a curated set of twenty to forty metrics, each with a name, a definition in plain language, an owner who is accountable for it, the precise filters and exclusions, the grain at which it is valid, and the synonyms people actually use when they ask for it. This is unglamorous, it takes a quarter, and it is the whole thing. Build it and the conversational layer becomes reliable. Skip it and no model, however capable, will rescue you. It also happens to be the least wasteful investment available right now, because it is vendor-independent. Whichever product wins, the certified definitions transfer.

Agree the metric before trusting the answerArticle-derived definition checklist, not a live ERP query, certified customer metric or measured implementation timeline.
DecisionWhat to make explicit
MeaningThe financial or operational basis of the metric
ScopeEntities, grain and reporting boundaries
ExclusionsCancelled, test, legacy and intercompany records
AuthorityThe asking user's own authorisations
EvidenceThe query, filters and accountable owner

Qualitative summary of this article's source text, not a measured outcome or performance estimate.

Where this genuinely works today

Lookup and navigation, not analysis. What is the status of this purchase order. Which invoices for this customer are more than thirty days overdue. When is this item next due in. What did we last pay this supplier for this part. These questions have one correct answer, the answer is a record rather than a computation, and there is nothing for a definition to be wrong about. That is not a small prize. Replacing six screens of navigation and a transaction code nobody remembers with one sentence is worth real money across a large user base, and it carries almost none of the risk that analytical questions carry. Deploy there first, and let the analytical use cases wait for the semantic layer.

How to run the evaluation

Refuse the vendor's dataset. Insist on yours, with your customisations and your company codes. Bring three questions your finance team already argues about. Ask what happens when a question is ambiguous, and treat a system that asks a clarifying question as better than one that answers smoothly. Ask to be shown a case where it gets the answer wrong, and be sceptical of any vendor who says there isn't one. Then test with an ordinary user's permissions rather than an administrator's.

Practical Guidance for AI-Powered ERP Query Implementation

  • Start with lookup questions, not analytical ones.
  • Build the certified metric set before buying the interface.
  • Require execution under the user's own authorisations.
  • Test with restricted accounts, not administrator accounts.
  • Insist on evaluation against your own data and your own disputed questions.
  • Prefer systems that ask clarifying questions over systems that always answer.
  • Show the query and the filters behind every number.
  • Name an owner for each metric definition.

The Regional Angle

Three conditions common to groups here make the definition problem sharper than it is elsewhere. The first arrived this year. Corporate tax in the United Arab Emirates applies to financial years beginning from June, zakat and tax operate alongside each other in Saudi Arabia on different bases, and free zone entities may sit under a separate treatment subject to qualifying conditions. A regional group therefore now holds at least three defensible answers to what profit was: the management figure the board discusses, the statutory figure in each entity's accounts, and the taxable figure computed under a regime whose implementation details are still being worked through. Ask a conversational layer what profit was last quarter and it will return one of them, without telling you which. Before deploying anything that answers financial questions in natural language, write down which basis is the default, label the others explicitly, and make the system state the basis in every answer rather than assume it. The second is structural. The operating business that people have in mind when they ask a question is rarely a single legal entity here: mainland companies, free zone entities, offshore holding vehicles and joint ventures with local partners, frequently a dozen or more for a mid-sized group, each with its own company code and some with their own charts of accounts. The question means the business; the data means the entities; and the gap between them is filled by a consolidation hierarchy plus intercompany eliminations that usually live in a spreadsheet maintained by the group financial controller. Until that hierarchy exists as data the ERP can read, every aggregate question is answered against a structure the asker did not intend, and the error will be large rather than marginal because intercompany traffic between group entities is often a substantial share of turnover. The third is political rather than technical, and it decides whether any of this gets used. In most regional finance functions the real reporting layer is a workbook owned by one person, reconciled by hand, trusted by the chief executive, and different from what the system produces for reasons that have accumulated over years. When the conversational layer disagrees with that workbook, the workbook wins, and the new tool is quietly abandoned as unreliable. The only way through is to run the two in parallel for a period, investigate every difference, and resolve each one explicitly — either the system definition is corrected or the workbook adjustment is documented as a deliberate rule and moved into the system. That reconciliation is the adoption programme. Skipping it guarantees the tool is used for lookups and ignored for anything that matters.

The objection worth taking seriously

The strongest objection is that this argument has been made against every reporting tool for thirty years, and the tools shipped anyway and proved useful. Self-service analytics was going to produce chaos; instead people learned which numbers to trust and which to check with finance. Demanding a complete certified semantic layer before anyone can ask a question in plain language is a counsel of perfection that will delay a genuinely useful capability by a year, during which staff will keep exporting to spreadsheets and computing far worse numbers by hand. Imperfect answers from a governed system beat unauditable answers from a laptop. This is largely right, and it is why the recommendation is to deploy for lookup immediately rather than wait. The difference from previous generations is where the definition lives. A dashboard embodies its definition in an artefact: someone built it, the logic is inspectable, and when it is wrong you can find out why and fix it once for everybody. A conversational answer embodies its definition in a query that existed for one second, was seen by one person, and left no record. When the board is given a number that turns out to be wrong, there is nothing to examine and nobody to ask. That is a genuinely new problem, and the fix is not perfection before launch — it is showing the query and the filters alongside every answer, so that the artefact exists.

Common Questions

Can we just point a model at the ERP database?

Technically yes, and several teams have. The results are unreliable for anything beyond simple lookups, and the permission model is usually bypassed in the process, which turns a productivity experiment into a disclosure incident.

Is this better on a modern cloud ERP than an older on-premise system?

Somewhat, because the data model is more accessible and the vendor's own tooling is available. The definition and permission problems are identical, and heavily customised cloud systems present the same difficulties as old ones.

How long does a semantic layer take?

For twenty to forty core metrics, roughly a quarter, and most of that time is spent getting people to agree rather than building anything. The agreement is the deliverable.

What should we expect over the next twelve months?

Expect every major ERP vendor to ship a conversational layer in its next release cycle, and expect most of those to be read-only and restricted to a narrow question set at first, for the permission reasons above. Expect the certified metric definition to emerge as the contested asset, with vendors offering to build it for you and consultancies selling it as a product. Expect at least one prominent case of a board receiving a confidently wrong figure from one of these tools, and expect it to set expectations for the whole category. And expect the honest benefit in the first year to be the disappearance of transaction codes, not the disappearance of the finance analyst.


AI-Powered ERP Query Implementation — we build the certified metric definitions first, wire the query layer to your real authorisation model, and start where the answers cannot be wrong.

Continue reading

Talk to OPS

Start with the operating problem.