Collaboration / Source date:

Documentation Written for AI Retrieval, Not Just Humans

Structure, headings, and explicit context improve both assistant accuracy and human comprehension.

Illustration of an editor reviewing a self-contained expense-policy excerpt with entity, country, date and owner fields.

Most organisations that have connected an assistant to their internal documentation this year have had the same experience. The demonstration was excellent. The rollout produced confident, plausible, wrong answers, and the instinct was to blame the model. It is usually not the model. It is the documentation, which was written on assumptions that retrieval does not honour.

Retrieval does not read your documentation. It reads a few hundred words of it, chosen without context, and answers as though that was the whole page

Everything below follows from that single mechanical fact.

What actually breaks

Pronouns and back-references. A passage beginning "this process applies only in the second case above" is meaningless once it is separated from the paragraph that defined the cases. It will still be retrieved, and it will still be answered from. Implicit scope. A policy document whose title says which country, entity or customer tier it covers, and whose body never repeats it, loses that qualification the moment a section is retrieved on its own. Stale duplicates. Three versions of the expenses policy exist: the current one, the 2021 one in a departmental folder, and a draft. A human picks the right one by looking at the folder. Retrieval has no folder intuition and will cheerfully quote the draft. Meaning carried by layout. Screenshots, annotated diagrams and wide tables where the sense lives in the column headers all degrade badly. If the answer is only visible to an eye, it is invisible to retrieval. Headings that hide the answer. A section called "Background" containing the actual rule will be retrieved less often than a section called "Approval limits" containing nothing useful.

Six rules that fix most of it

Write each section so it stands alone, repeating the scope in the first sentence. State the entity, country and effective date inside the text, not only in the file properties. Put the answer in the heading language people would actually use when asking. Convert screenshots and complex tables into prose or simple lists where the content matters. Keep one topic per page, so the retrieved fragment and the page agree. And record the owner and the last review date in the body. None of that is exotic. It is the plain-language editing most internal documentation has never received.

The real work is deletion

The single biggest improvement in answer quality in most deployments comes from removing documents, not adding them. Superseded policies, abandoned project spaces, duplicated handbooks and old templates are not inert; they compete with the current version and sometimes win. This is unpopular work because nobody wants to be the person who deleted something. Make it easier by archiving to a location the assistant cannot see rather than destroying anything. The retrieval boundary, not the recycle bin, is the control you actually need.

Metadata earns its keep only if it travels

Three fields matter: who owns this, when does it take effect, and what does it apply to. They matter because they can be used to filter retrieval and to warn a reader. They only work if they are inside the content as well as attached to it, because many pipelines index the text and drop everything else.

Test it like a support queue

Collect the questions people actually ask, ask them of the assistant, and for each wrong answer look at which passage was retrieved. Nearly every failure resolves to one of the five causes above, and the fix is an edit to a document rather than a change to the system. Keep the question list and re-run it after each cleanup round, because that is the only evidence you will have that the work paid for itself.

Practical Guidance for Documentation Structure Review

  • Make every section self-contained, with scope restated in its first line.
  • Put the entity, country and effective date inside the text.
  • Title sections with the question people ask, not the topic area.
  • Replace screenshots and wide tables where they carry the answer.
  • Archive superseded documents outside the retrieval boundary.
  • Name an owner and a review date in the body of each page.
  • Keep one answer per page so fragments and pages agree.
  • Re-run a fixed question set after every cleanup round.

The Regional Angle

The most common retrieval failure in this market comes from a document structure that looks entirely sensible on screen: a single group policy — leave, end of service, expenses, procurement authority — with a section for each country of operation. A human scrolls to their heading. Retrieval returns whichever section matched the wording of the question, which is frequently the wrong jurisdiction, and the answer arrives with no indication that a different rule applies to the person asking. Given how sharply entitlements diverge across the Gulf and North Africa, that is not an inconvenience; it is a wrong entitlement quoted to an employee in writing. Split multi-jurisdiction policies into one document per entity, put the country in the title and in the first line of every section, and keep the comparison table as a separate reference page that says explicitly that it is not authoritative. The second is the part of the operation that has never been written down at all. Government relations, licence renewals, visa and labour processes, customs clearance, chamber attestations and municipality approvals are typically held in the heads of long-serving public relations officers and administrators who know which portal behaves oddly, which document the counter will reject, and what the real sequence is this month. An assistant trained on your document store amplifies the well-documented half of the business and leaves that dependency exactly where it was — arguably worse, because the polished answers create an impression of coverage. Pick the three government-facing processes whose interruption would hurt most, sit with the person who runs them, and write the current sequence down with the failure points named. It is the highest-value documentation work available in a regional group and it has nothing to do with artificial intelligence. The third is format. A great deal of regional internal governance circulates as signed and stamped PDF memoranda, scanned circulars and photographs of notices shared through messaging groups, because the signature and stamp are what give the instruction authority. Those files are frequently images with no text layer, which means retrieval cannot see them at all, and the assistant will answer from the older typed policy they were issued to override. Run recognition over the circular archive so the text is extractable, and maintain a simple register of circulars — number, date, what it changes, whether it is still live — as a text page. The stamped original stays the authority; the register is what makes it findable.

The objection worth taking seriously

The strongest objection is that this inverts the purpose. Documentation exists for people, and rewriting it to suit the ingestion characteristics of a retrieval pipeline is a concession to a tool that will be obsolete in two years. Models with far larger context windows are already appearing; before long the system will simply read the whole document and the careful chunk-friendly restructuring will have been an expensive accommodation to a temporary limitation. Meanwhile the repetition of scope in every section makes documents more tedious for the humans who still have to read them. The technical half of that is likely right, and anyone rebuilding a documentation estate purely around today's chunk sizes is making a mistake. What survives the technology change is almost everything on the list. Self-contained sections, stated scope, honest effective dates, headings that name the question, and the removal of three contradictory versions of the same policy are improvements for a new joiner, an auditor and a colleague in another country, entirely independent of whether a machine reads them. The genuinely retrieval-specific item is the deletion work, and larger context windows make that worse rather than better: a model given the whole estate now sees all three contradictory expense policies at once. The correct summary is that assistants did not create a documentation problem. They made an existing one legible, and they made it embarrassing.

Common Questions

Should we rewrite everything before deploying?

No. Deploy against a narrow, high-quality corpus, then expand. A small clean set answers better than a large messy one, and it gives you a standard to hold new content to.

Does adding more documents improve answers?

Only if they are current and non-contradictory. Volume without curation reliably reduces accuracy, because the competing passage only has to win once.

Who should own this work?

Someone with authority to delete. Documentation cleanup fails when the owner can only request changes from the teams that wrote the originals.

What should we expect over the next twelve months?

Expect context windows and retrieval quality to keep improving, and expect that to raise the floor rather than remove the need for curation. Expect platforms to add freshness and authority signals so recent, owned documents outrank abandoned ones. Expect documentation ownership to appear in job descriptions where it never has before. And expect the organisations that spent this autumn deleting to be the ones whose assistants are trusted next year.


Documentation Structure Review — we find the passages your assistant is answering from, and fix the documents rather than the prompt.

Continue reading

Talk to OPS

Start with the operating problem.