For twenty years the answer to organisational forgetting has been the same: build a place to put things. An intranet, then a wiki, then a knowledge base, then a documentation site, each launched with a taxonomy workshop and abandoned within eighteen months to the people who like writing documentation. This year the vendors changed the proposition. Instead of a new place to store knowledge, they are offering retrieval over the content you already have — the shared drive, the ticket history, the chat archive, the contracts folder. The question is no longer where knowledge should live. It is whether anything in the places it already lives can be found.
Twenty years of knowledge management asked people to store things properly. The only question that ever mattered was whether anyone could find them afterwards
The storage model failed for a reason that was never a tooling problem. Writing something down properly costs the author an hour and benefits a stranger six months later. That trade is unattractive in every organisation on earth, and no amount of platform quality changes it. Maintenance is worse: nobody is paid to revisit a document that is quietly rotting, so the knowledge base fills with material that was true in 2019. And the moment you copy an answer into a knowledge base, it forks from the source. The original gets updated. The copy does not. Now there are two answers and no indication which is current. Retrieval over existing content removes the copy and the forced migration. It does not remove the rot.
Retrieval quality is bounded by the corpus, and the corpus is the problem
A system that searches semantically, ranks well and synthesises a clean answer is still reading your actual documents. If there are eleven versions of the travel policy in four folders, it will find one of them and present it with total confidence and a citation, which is more persuasive than the old behaviour of returning eleven links and letting the reader notice the contradiction. Confident synthesis over a contradictory corpus is worse than a search box, because the contradiction is now hidden.
Three things that actually determine whether this works
Authority. The system needs some way to know which document wins. An owner, a review date, an explicit current marker, a canonical location. Most corpora have none of these, which means ranking falls back on textual similarity — and the obsolete document often matches the question better, because it was written when people used those words. Deletion. The highest-value knowledge activity in a retrieval world is archiving. Every superseded policy, abandoned project plan and draft that was never approved is now an active liability rather than harmless clutter. Before buying anything, run an archive campaign. It is unglamorous, it is fast, and it improves answer quality more than any feature. Structure. Long narrative documents retrieve badly because the relevant passage is buried in context the system has to guess at. Documents with clear headings, question-shaped titles and sections that make sense read alone retrieve well. Writing style has quietly become an infrastructure decision, and the practical instruction to your teams is simple: write the heading as the question someone will ask.
You cannot evaluate this on a vendor demo
Build a question set. Take fifty real questions from your help desk queue, your internal chat channels and the things new joiners ask in their first month. Write down the correct answer for each, verified by whoever owns the subject. Then run every candidate system against the same list and score it. This takes two days and it is the only procurement input that means anything. Every product demonstrates beautifully against its own curated corpus. The interesting number is how many of your fifty it gets right, and more importantly how many it gets confidently wrong, because those are the ones that will cause damage.
Measure deflection, not usage
Usage statistics for these tools are meaningless and every vendor will offer them. The questions worth tracking are whether the number of repeat questions reaching your support and HR teams falls, whether time to first correct answer drops for new joiners, and what proportion of answers people actually act on without checking with a colleague. If the volume of questions to the shared inbox has not moved after a quarter, the tool is entertainment.
The ninety days before you buy
Assemble the question set. Put an owner and a review date on your two hundred most-read documents. Archive everything superseded, ruthlessly. Consolidate the duplicate policies down to one of each. Then run the procurement. Organisations that do this first get a working system in a quarter. Organisations that buy first spend a year explaining why the answers are wrong.
Practical Guidance for Knowledge Retrieval Assessment
- Build a fifty-question test set from real questions with verified answers.
- Archive superseded content before connecting anything to it.
- Assign owners and review dates to your most-read documents.
- Consolidate duplicate policies to a single canonical version.
- Write headings as the questions people ask.
- Score vendors on confident wrong answers, not just correct ones.
- Track deflection from support queues, not tool usage.
- Check every answer shows its source document and date.
The Regional Angle
Three conditions here change what corpus preparation means in practice. The first is ownership decay at a rate most head offices do not anticipate. Workforce mobility in the Gulf is high, departures are often abrupt because they are tied to visa and contract events rather than notice periods, and people leave the country rather than moving to a competitor down the road. The result is a document estate in which a large proportion of content has an owner who is no longer reachable by any means. When the retrieval layer surfaces a procedure and someone asks whether it is still correct, there is nobody to ask. Run an ownership validation sweep against the active employee list before you connect anything: reassign what matters, archive what nobody will claim, and add document ownership transfer to the offboarding checklist alongside the laptop and the visa cancellation. It is the single highest-yield hour of corpus work available here. The second is that policy content is jurisdictionally plural in a way that makes the most commonly asked questions the most dangerous ones. Annual leave, notice periods, end-of-service gratuity, probation, overtime and medical cover differ between a mainland company and a free zone entity, differ again in Saudi Arabia, and differ again in Egypt or Jordan where the group's back office may sit. Staff ask these questions constantly, and a retrieval system that returns the gratuity calculation for the wrong entity gives an answer that is not merely unhelpful but financially wrong and quotable back at you. Jurisdiction and entity have to be metadata on every HR and finance policy document, and the retrieval layer has to filter on the asker's own entity rather than rank across all of them. If your candidate system cannot do that filtering, it is not deployable for the highest-volume question category you have. The third concerns how questions are typed rather than how documents are written. Most of the workforce here is asking in a second or third language, and real queries are short, keyword-shaped, inconsistently spelled and frequently mixed between languages. Systems tuned on fluent, well-formed English questions degrade sharply against that input, and the degradation is invisible in a demo because the demonstrator types perfect sentences. When you build the fifty-question test set, harvest the questions exactly as people actually wrote them, misspellings included, rather than tidying them into good English first. The difference in scores between the tidy version and the real version is the number you should be procuring against.
The objection worth taking seriously
The strongest objection is that this is enterprise search wearing a new hat. Organisations bought federated search in the 2000s, bought it again as intranet search in the 2010s, and each time the pattern was identical: impressive pilot, mediocre production, quiet decline. The reason was never the ranking algorithm. It was that the underlying content was contradictory, stale and incomplete, and no retrieval mechanism repairs a corpus. If the documents are wrong, a better way of finding them produces wrong answers faster. That history is accurate and the corpus point stands, which is why the recommendation here is to spend the first quarter on content rather than on procurement. Two things have genuinely changed, though, and they are worth being precise about. Semantic matching removes the requirement that the asker guess the words the author used, which was the failure mode that killed keyword search for people who did not already know the jargon — which is to say, exactly the people who needed to search. And synthesis removes the reading cost: the previous generation returned twelve documents and the user opened none of them, because finding the answer still required twenty minutes. Those two changes are real and they are why the category is worth revisiting. They just do not substitute for deleting the eleven old travel policies.
Common Questions
Should we still maintain a wiki?
Yes, for the small set of material that has no natural home elsewhere and genuinely needs authoring — onboarding paths, decision records, how-we-work documents. Stop using it as a copy of content that lives somewhere authoritative.
Will this work over chat history?
Partially. Chat contains a great deal of true, current knowledge and an equal quantity of speculation, jokes and superseded plans, with no signal distinguishing them. Index it only if the system can weight it below authored documents.
How do we stop it answering from old documents?
Archive them. Recency weighting and review dates help at the margin, but the reliable control is that the obsolete document is not in the index at all.
What should we expect over the next twelve months?
Expect retrieval features to appear inside every major collaboration suite by the middle of next year, bundled into licences you already hold rather than sold separately, which will make standalone products in this category a difficult business. Expect the competitive differentiator to shift from answer quality to permission handling and freshness, because those are the parts that break in production. Expect a run of internal incidents in which a confident answer drawn from a superseded document reaches a customer or a regulator, and expect those to do more for archiving discipline than a decade of governance policy. And expect the knowledge management role, where it still exists, to be redefined around curation and deletion rather than authoring — which is what it should always have been.
Knowledge Retrieval Assessment — we build the question set from what your people actually ask, clean the corpus that will answer them, and score the products against your content rather than the vendor's.
