Cybersecurity / Source date:

SIEM Deployments That Generated Alerts Nobody Read

Log collection without tuning or staffing produced compliance evidence instead of detection capability.

Illustration of a technician tracing a single network lead, representing investigation amid noisy security telemetry.

Security information and event management was, by 2010, the most reliable line item in an enterprise security budget. Compliance regimes required log collection and review. Auditors asked to see it. Vendors had products that collected everything and correlated it into alerts. Boards approved the spend without argument. And then, in organization after organization, the deployment finished, the dashboard lit up, and nobody looked at it again. The pattern was consistent enough to become an industry joke. Companies spent heavily on detection capability, produced tens of thousands of alerts a day, staffed a team that could realistically investigate a few dozen, and discovered during their next incident that the evidence had been sitting in the console for weeks.

How the Failure Actually Happens

It is tempting to blame the tool, but the failure is architectural and it repeats regardless of product. As SANS has argued more recently, SIEM struggles are usually symptoms of misalignment between technology, team and priorities rather than a technical defect.[1] The sequence is predictable. Everything gets connected because coverage is the procurement metric. Firewalls, servers, endpoints, applications, network devices, databases. The project is measured by log sources onboarded, so log sources are onboarded. Default rules are enabled because nobody wants to miss anything. Out-of-the-box correlation content is designed to demonstrate breadth in a proof of concept, not to fit one environment. It fires constantly. The noise is accepted rather than tuned. Tuning requires understanding which activity is normal in this specific environment, which requires time the deployment project does not budget for, because the project ends at go-live. Analysts adapt by triaging the categories they trust. They learn which alert types are always false positives and stop opening them. This is rational behaviour and it is where detection dies — desensitisation to a high volume of alerts is the textbook definition of alert fatigue.[2] The genuine alert arrives in a category that has been mentally filed as noise. Nobody opens it. The post-incident review finds it, timestamped, weeks earlier.

The Measurement Problem Underneath

The deeper reason this persisted for a decade is that the metrics reported to management measured the wrong thing. Log sources connected. Events per second. Alerts generated. Storage retained. Every one of those is a measure of collection, not of detection. An organization can score perfectly on all of them while being incapable of noticing an intrusion. The measures that would have exposed the problem were rarely reported: what proportion of alerts were actually investigated, how long triage took, what the false positive rate was per rule, which specific attack behaviours the rule set could detect, and whether any real incident had ever been found by the platform rather than by a user complaint or an external notification. That last question is the uncomfortable one, and it remained uncomfortable for years.

Measure detection rather than collectionArticle-derived review questions. No alert volumes, detection rates or performance improvements are claimed.
Collection measureDetection question
Log sources connectedWhich relevant attack behaviours can be detected?
Alerts generatedWhich alerts were investigated and how quickly?
Events per secondWhich rules create false positives?
Storage retainedDid a scenario reach a human and produce a response?

Qualitative summary of this article's source text, not a measured outcome or performance estimate.

Building Detection That Works

  • Start from threats, not from log sources. Decide which attack behaviours matter for your business — credential abuse, privilege escalation, bulk data export, unusual administrative activity — then collect the logs that reveal those specifically. Coverage-first deployments produce volume; threat-first deployments produce detection.
  • Size the rule set to your triage capacity. If you can investigate forty alerts a day, a configuration producing four hundred is not more secure, it is less. Fewer, higher-fidelity rules beat comprehensive noise every time.
  • Tune continuously and own it. Log source relevance should be reviewed regularly, with low-value noisy sources excluded rather than tolerated.[3] Tuning is an operating activity, not a deployment phase.
  • Track false positive rate per rule. Any rule above a defined threshold gets fixed or disabled. A rule nobody trusts is worse than no rule, because it occupies attention.
  • Write the response before the rule. Every alert needs a documented action. If nobody can say what to do when it fires, it should not fire.
  • Involve responders in detection design. Detection engineers building rules without the people who have to act on them is a standard cause of unusable output.[1]
  • Monitor the monitoring. A log source that silently stops sending is an invisible blind spot. Alert on collection gaps as seriously as on security events.
  • Test with real scenarios. Run a controlled simulation of the behaviour you claim to detect and confirm that an alert fires, reaches a human, and produces a response. Most organizations that do this for the first time are unpleasantly surprised.

The Compliance Trap

A large share of SIEM deployments in this period existed because a regulation or standard required log collection and review. That origin shaped the outcome. When the objective is passing an audit, the deliverable is evidence that logs are collected and retained. Nothing in that objective requires detection to work. So the project optimises for demonstrable coverage, the auditor is satisfied, and the security outcome is incidental. This is not an argument against compliance-driven investment — it funded capability that would otherwise have gone unfunded. It is an argument for being clear internally about which objective a given control serves. A platform built to satisfy an auditor and a platform built to catch intrusions look similar on a procurement document and behave completely differently at three in the morning.

Where This Goes Next

The volume problem has only intensified. Cloud platforms, containers, SaaS applications and identity providers all generate telemetry, and the total is orders of magnitude beyond what any analyst population can review. The current answer is automation and AI-assisted triage — correlating, enriching and closing low-risk alerts without human involvement, so that analysts see a smaller set of higher-confidence cases. Where this is implemented carefully it genuinely helps, and it is the only response that scales with the telemetry. But it recreates the 2010 failure in a new form if the same discipline is missing. An automated triage layer that closes alerts nobody audits is indistinguishable from an analyst who has stopped opening a category, except that it operates faster and produces a cleaner-looking dashboard. The control that matters is sampling: independently reviewing a proportion of what the automation closed, and measuring how often it was wrong. The lesson from a decade of unread alerts is not that detection technology fails. It is that detection is an operating discipline with a technology component, and that the organizations which measured collection instead of investigation spent a great deal of money learning it the hard way.

Common Questions

What is SIEM alert fatigue?

The condition where analysts become desensitised to a high volume of alerts and stop investigating whole categories, typically because default correlation rules were deployed without environment-specific tuning.

Why do SIEM deployments fail to detect real incidents?

Because they are built around log source coverage rather than specific threat behaviours, generate far more alerts than the team can triage, and are measured on collection metrics that say nothing about whether anything is being investigated.

What metrics indicate a detection programme is actually working?

Proportion of alerts investigated, mean time to triage, false positive rate per rule, documented coverage of specific attack behaviours, and whether real incidents have been discovered by the platform rather than reported externally.

Does AI-assisted triage solve alert overload?

It is the only approach that scales with modern telemetry volumes, but it reproduces the original failure unless a sample of automatically closed alerts is independently reviewed and the error rate measured.


Detection Capability Review — Outpace tests whether your monitoring would actually catch the attacks that matter, tunes the noise out, and replaces collection metrics with measures of detection that mean something.

Continue reading

Talk to OPS

Start with the operating problem.