Collaboration / Source date:

Intranet Search Was the Real Knowledge Problem

Content existed but could not be found, so employees recreated documents rather than search for them.

Illustration of a knowledge coordinator retrieving the current approved policy for a colleague.

Every knowledge management programme of the 2000s ended at the same place. The organization built a portal, migrated documents into it, wrote a taxonomy, appointed content owners, and launched with an internal communications campaign. Within eighteen months, employees had gone back to emailing colleagues to ask where things were. The usual diagnosis was cultural: people did not contribute, teams hoarded knowledge, nobody maintained their sections. That explanation was comfortable because it made the failure someone else's fault. The more accurate diagnosis was mechanical. The content usually existed. People could not find it. Research on enterprise search implementation has consistently examined barriers in exactly these terms — how results are presented, how people search, and why organizational portals fail to return what users need.[1]

By 2010, employees had a decade of experience with web search that worked. They typed a few words and got the right answer. Then they arrived at work, typed the same kind of query into the intranet, and received eleven hundred results sorted by modification date. The gap had structural causes. There were no links to rank by. Web search ranking was built on the link graph — pages that many other pages point to are probably important. Internal documents do not link to each other. Without that signal, relevance ranking had almost nothing to work with beyond keyword frequency. Content had no quality signal. A superseded draft from 2006 and the current approved policy looked identical to the index. Nothing recorded which one people actually used, whether it had been reviewed, or whether its owner still worked there. Vocabulary did not match. Employees searched for what they called things; documents were titled with what the authoring department called them. Someone looking for "holiday allowance" did not find "Annual Leave Entitlement Policy v3 FINAL." The content was scattered across systems search could not reach. File shares, the document management system, the intranet, mailboxes, ticket histories and departmental databases. Even excellent search over one repository returns a fraction of the answer when the organization's knowledge is distributed across six. Permissions made indexing hard. Search that ignores access control leaks information; search that respects it must evaluate permissions per user per result, which in 2010 was slow and frequently implemented by simply not indexing sensitive repositories — which happened to be where the valuable material lived.

The Consequence Nobody Measured

The cost of bad internal search is almost entirely invisible because it shows up as other things. It shows up as work being redone, because the previous version could not be found. As inconsistent answers to customers, because three people found three different documents. As new joiners taking longer to become productive. As experienced staff being interrupted constantly, because asking a person is faster than searching. And as compliance exposure, because the current policy exists but the version people actually use is four years old. None of that appears on a line item. It appears as general organizational friction, which is why knowledge management budgets were persistently cut and persistently reinstated.

Findability needs more than matching wordsQualitative search problems and responses drawn from the article, not a measured search benchmark.
Search problemReview focus
Current and superseded documents look alikeNamed owners, review dates and explicit authority
Staff and authors use different wordsSearch logs and reader vocabulary
Knowledge is scatteredRepositories the search index can actually reach
Results reveal restricted contentAccess controls on titles, snippets and answers

Qualitative summary of this article's source text, not a measured outcome or performance estimate.

What Actually Improved Findability

The organizations that made progress in this era did a small number of unglamorous things.

  • Fix the small set of high-traffic answers first. Analyse what people actually search for. A surprisingly small number of queries — leave policy, expense limits, who approves what, how to raise a purchase order — account for a large share of volume. Curating those properly delivers more value than a full taxonomy.
  • Use search logs as a requirements document. Queries that return nothing, or where nobody clicks a result, tell you precisely what content is missing or badly titled. Almost nobody reviewed these logs; it is the cheapest improvement available.
  • Title documents the way people speak. A title matching the user's vocabulary beats a correct formal name. Add the informal terms as keywords rather than insisting on institutional language.
  • Give content a lifecycle. Review dates, owners, and automatic archiving of material nobody has opened in two years. Unmanaged accumulation is the main cause of poor relevance, because the index fills with material that should not be there.
  • Mark authority explicitly. Which document is the current approved version, who owns it, when it was last reviewed. Both search ranking and human confidence depend on this, and it is usually absent.
  • Index across systems, not just the portal. Federated or unified search across the repositories people actually use is worth more than improving any single one.
  • Respect permissions in results, not after them. Users should see what they can access, without exposure through titles or snippets.
  • Measure findability, not contribution. Successful searches, zero-result rate, and time to answer. Counting documents uploaded measures activity; counting answers found measures the outcome.

Why the Problem Persists

Fifteen years later, this remains substantially unsolved, and the people running the systems say so openly. Content and knowledge now live across even more places — HR and service management systems, intranets, ticket histories, file repositories and chat platforms — which is exactly why employees still ask for direct links in Slack or Teams rather than searching.[2] The fragmentation got worse, not better. Every SaaS application added its own repository with its own search box, and the average organization now holds its institutional knowledge across dozens of systems rather than six.

What AI Changes and What It Does Not

Retrieval-augmented AI assistants are the most significant improvement to this problem in twenty years, and the improvement is real. Semantic matching solves the vocabulary mismatch — asking about "holiday allowance" now finds the annual leave policy. Synthesis across multiple documents answers questions no single document answers. And a conversational interface removes the need to guess keywords. What it does not change is the content problem underneath, and this is where organizations are currently making an expensive mistake. An assistant retrieving from a repository containing four versions of a policy will answer from one of them, fluently and without hesitation. The stale draft and the approved current version look the same to a retrieval system unless something distinguishes them. The failure mode is worse than bad search results, because bad search results are visibly bad — the user sees eleven hundred hits and knows to be careful. A confident synthesised answer from a superseded document carries no such warning. Which means the boring work that was skipped for twenty years has become a prerequisite. Content lifecycle, ownership, review dates, explicit marking of the authoritative version, and archiving of material that should no longer be consulted. Those were good hygiene in 2010. With an AI layer on top, they are the difference between an assistant that is trustworthy and one that is confidently wrong at scale. The knowledge problem was never really about tools. It was about whether anyone owned the currency and authority of the content — and that is still the question.

Common Questions

Internal documents lack the link graph that powers web ranking, there is no quality or popularity signal to distinguish current from superseded content, user vocabulary rarely matches document titles, and content sits across multiple systems that search cannot reach uniformly.

What is the real cost of poor internal findability?

Work redone because prior versions cannot be found, inconsistent answers to customers, slower onboarding, constant interruption of experienced staff, and compliance exposure when people use outdated policy documents.

What improves enterprise search most cheaply?

Reviewing search logs for zero-result and no-click queries, curating the small number of high-traffic answers, titling content in users' vocabulary, and archiving material nobody has opened in years.

Does AI-powered search solve the knowledge management problem?

It solves vocabulary mismatch and synthesis across documents, but not content currency. An assistant will answer confidently from a superseded document, which makes content lifecycle, ownership and explicit authority marking more important than before, not less.


Enterprise Search Assessment — Outpace finds out what your people actually search for and fail to find, cleans up the content that makes every answer unreliable, and gets your knowledge base ready for an AI layer that will otherwise repeat your worst documents back to you.

Continue reading

Talk to OPS

Start with the operating problem.