Knowledge management was always a discipline with a compliance problem. People wrote documentation because they were asked to, other people did not read it, and the value of the exercise was mostly invisible. What has changed is that the documentation now has a second reader — the retrieval layer sitting behind every assistant in the company — and that reader is both tireless and completely literal. That reader has different requirements from the human one, and organisations that never got documentation right are now discovering that the cost of not getting it right has gone up.
Documentation used to fail quietly. Now it fails out loud, in an answer given confidently to someone who had no way to know the source was three years out of date
Here is how to structure knowledge when half your readers are machines.
What the retrieval layer needs that humans did not
Self-contained units. A human reading a page understands it in the context of the section it sits in. A retrieved fragment arrives alone. Every meaningful chunk needs enough context to stand by itself, which usually means more repetition than good human writing tolerates. Explicit scope. "This applies to the UAE entity only, from January 2026" needs to be in the text, not implied by which folder the page is in. Folder structure is invisible to retrieval. Unambiguous currency. A last-edited timestamp is not a statement of validity. Content that is current should say so, and content that has been superseded should say what superseded it. One authoritative version. Humans cope with three copies by asking someone. Retrieval returns whichever copy scores highest, which may be the one in a 2023 project folder.
The two things that matter most
Deletion, which nobody does. Every obsolete document remaining in the corpus is an active source of wrong answers, and the single highest-value knowledge exercise available to most organisations is not writing but removing. Superseded content should be deleted or clearly marked as historical, and this needs to be somebody's standing responsibility. And the tacit layer. A large share of what an organisation knows has never been written down because it was cheaper to ask someone. Retrieval cannot reach it, so the gaps in the corpus are exactly the questions people most often ask. Work out what is being asked and not found — the query log is the best documentation backlog you will ever get.
What to stop doing
Stop measuring documentation by volume. Stop maintaining parallel structures for different audiences. And stop treating the wiki as an archive; an archive that is indexed is not an archive, it is a source.
State scope and validity
Put the applicable entity, period and authoritative version in the text.
Remove stale sources
Delete or mark superseded material and keep archives outside the retrieval index.
Find missing answers
Use failed queries to locate unwritten knowledge.
Assign ongoing ownership
Give a named custodian time to prune and review the corpus.
Qualitative summary of this article's source text, not a measured outcome or performance estimate.
Practical Guidance for Knowledge Architecture Review
- Make each section self-contained, with scope stated in the text.
- Delete or mark superseded content; this is the highest-value task.
- Establish one authoritative copy per subject.
- State validity explicitly, not via timestamps.
- Mine the query log for what people cannot find.
- Write down the tacit knowledge people ask each other for.
- Exclude archives from the retrieval index.
- Give the corpus an owner with time allocated to pruning.
The Regional Angle
The first regional issue is bilingual content, which creates a structural retrieval problem rather than an inconvenience. Policies exist in Arabic and English, sometimes with substantive differences, and a retrieval layer will happily answer an English question from an Arabic source or vice versa without flagging that the two versions diverge. Organisations should designate which language version is authoritative for each document type and record it in the text, because the alternative is confident answers drawn from whichever translation happened to rank. The second concerns the regulatory dimension of stale content, which is sharper here than in slower-moving jurisdictions. Gulf tax, labour and data protection rules have changed repeatedly over the past few years, and corpora in regional companies are full of guidance written before the last three changes. A retrieval system that answers a corporate tax question from a 2022 document is not merely unhelpful; it is a compliance risk, and this argues for treating regulatory content as a separately governed subset with explicit review dates. The third is about the tacit layer under regional workforce mobility, where the ordinary argument for documentation becomes an operational necessity. Where a meaningful portion of the team may leave the country within a few years, the knowledge that was never written down does not degrade slowly — it leaves on a flight. Retrieval gives that documentation effort a payoff that is visible immediately rather than at the point of departure, which is the first genuinely persuasive argument for it that most regional managers have ever had.
The objection worth taking seriously
The strongest objection is that this is asking people to write for machines, and that the result will be worse for the humans who still have to read it. Self-contained chunks with repeated context and explicit scope statements make for tedious prose, and the discipline described here is essentially search engine optimisation applied to internal documentation — an exercise that degraded the public web considerably. Organisations will end up with a corpus that retrieves beautifully and that no person can bear to read. There is something to this, and the comparison to what optimisation did to the web is a fair warning. But the two audiences want more of the same thing than the objection allows. Explicit scope, a clear statement of what is current, and one authoritative copy are improvements for human readers too — the person who reads a policy and cannot tell which entity or year it applies to is failed by the same defect that confuses retrieval. The genuinely machine-specific requirement is self-contained chunking, which costs some elegance, and that cost is small compared with the alternative. And the most valuable item on the list, deletion, makes the corpus strictly better for everyone. The risk of writing unreadable documentation is real but it is a failure of execution, not a property of the requirement.
Common Questions
Should we restructure everything before deploying retrieval?
No. Deploy, watch what it answers badly, and fix those areas. The failures tell you where the corpus is weak far more efficiently than an audit will.
How do we handle documents that are historically important but obsolete?
Move them outside the index. Retained is not the same as retrievable, and conflating the two is the most common cause of wrong answers.
Who should own the corpus?
Someone with time allocated to pruning rather than a committee that approves contributions. The bottleneck is removal, not addition.
What should we expect over the next twelve months?
Expect retrieval quality to become a visible measure of documentation quality. Expect regulatory content to be governed separately. Expect the query log to become a standard input to knowledge planning. And expect deletion to finally get resourced, for the first time, because it now has an observable cost.
Knowledge Architecture Review — we start with what to delete, because in a retrieval world the obsolete document is the one doing the damage.
