Conference season is over and the marketing has settled into a pattern. Every detection product now has an assistant, every vendor briefing includes the phrase "AI-powered", and every threat intelligence report warns that attackers are using the same technology against you. Both statements are true. Neither tells a security leader what to do differently this quarter. What has actually happened, eighteen months into broad availability of capable language models, is more specific and less dramatic than either side's messaging.
Both sides automated the cheap part of their work. Neither automated the judgement, and that is where your security programme now queues
That is the stalemate. It is not a stable equilibrium and it is not a crisis; it is a shift in where the bottleneck sits.
What defenders genuinely got
Three things, and they are worth separating because vendors sell them as one. Triage assistance. Summarising an alert, correlating it with related events, drafting the initial investigation narrative. This is real, measurable and the least exciting thing in any demonstration. It is also the only one with a defensible return this year. Natural-language querying of telemetry. For a team of three people who never learned the query language, this is transformative. For a team with a competent detection engineer, it is a convenience. Behavioural detection. Improved, incrementally, and mostly through better modelling rather than through language models. Vendors have been doing statistical anomaly detection for a decade; the branding changed faster than the capability. What defenders did not get is autonomous response worth trusting with anything consequential. Products offer it. Very few organisations have it enabled beyond isolating a laptop.
What attackers genuinely got
Also three things, and the same discipline applies. Better language, which removes the crudeness that used to function as a free control. Faster reconnaissance, because assembling a picture of an organisation from public sources is exactly the kind of tedious synthesis these tools do well. And personalisation at volume, which collapses the old distinction between the mass campaign you could filter and the targeted approach you had to think about. What they did not get is new exploit classes, novel malware families that defeat endpoint detection, or the ability to skip the part where they need valid credentials or an unpatched service. The gain is throughput and polish. That matters, but it is a different claim from the one the more excitable reports make.
Why this produces a queue rather than a winner
Defensive automation increases the number of alerts that get written up competently. Offensive automation increases the number of attempts that look plausible enough to escalate. Both sides compressed the labour cost of volume, and on both sides the remaining constraint is a person deciding something. The practical consequence in most organisations is that the same number of analysts now face a longer queue of better-documented items, while the proportion of genuinely serious events in that queue falls. That is a worse working environment and roughly unchanged risk, which is a hard thing to put in a board paper after a year of investment.
The measurements that survived contact
Alert volume is now actively misleading, because both the tooling and the attackers inflate it. Three numbers hold up better. Time to decision per escalated alert — not time to close, which auto-closure flatters. The proportion of alerts no human ever reviewed, which is the number that tells you where your real exposure now sits. And, if your tooling supports it, the rate at which auto-closed items were later reopened, which is the only empirical check on whether the automation is calibrated or just fast.
Four questions for any vendor claiming AI defence
What was the base rate before the model, measured on our kind of environment rather than on a benchmark. What happens operationally when it is wrong, and who finds out. Does it act or does it advise, and can that be configured per detection class. And can we see the evaluation set, because a product evaluated against yesterday's attacks is a product that will surprise you. An answer of "the model learns from your environment" to the first question is not an answer.
Where the money should probably go instead
The overwhelming majority of successful intrusions still begin with valid credentials or an unpatched internet-facing service. Language models changed neither. A budget cycle spent on phishing-resistant authentication, privileged access hygiene and a reliable asset inventory buys more risk reduction than the same money spent on an assistant that writes better tickets. This is unfashionable advice in June 2024, and it has been correct every year for a decade.
Practical Guidance for AI Security Strategy Assessment
- Separate triage gains from detection claims when evaluating products.
- Measure time to decision, not alert volume or time to close.
- Track what proportion of alerts no human sees, and review a sample.
- Keep autonomous response scoped to reversible actions.
- Ask vendors for the base rate and the evaluation set in writing.
- Fund identity and patching before assistants.
- Retire language-quality cues from awareness training.
- Re-run one tabletop with a convincingly written lure.
The Regional Angle
The most consequential local change is that the quality of Arabic in attack content crossed a threshold this year. For a long time, poorly translated Arabic phishing was an informal control in this market — staff spotted the stilted phrasing, the wrong register, the mixed dialect, and forwarded it to information technology. That control has quietly expired. Machine-translated Arabic is now idiomatic enough to pass, including in the formal register used for official correspondence, and convincing enough in Gulf dialect to work in messaging apps. Two consequences follow. First, awareness training material in most regional organisations is written in English and teaches language tells, which means it is now training people on a signal that no longer exists. Second, the tells that remain are structural rather than linguistic — an unexpected change of payment details, a request routed outside its normal process, urgency attached to a transaction. Rewrite the training in Arabic as well as English, and rewrite it around process anomalies rather than spelling. The second is how managed detection is contracted here, which deserves scrutiny this year specifically. A large share of regional organisations buy monitoring from a provider priced on devices or log volume, reporting monthly on alerts raised and incidents handled. As both defensive tooling and offensive automation inflate alert counts, that report improves while your actual position does not — the provider's headline metric and your risk have become decoupled. Change the reporting before the next renewal. Ask for time to containment, the number of alerts closed without human review, the reopen rate on those, and a named analyst who can explain a specific decision. A provider who cannot produce those numbers is selling you volume, and volume is the one thing this year has made abundant. The third is the shape of the regional threat picture, which distorts the dashboards. Activity here spikes around geopolitical events, and those spikes are overwhelmingly high-volume, low-sophistication campaigns — defacement, credential stuffing, opportunistic scanning, claimed breaches that turn out to be recycled data. Automated triage handles that traffic well, which is genuinely useful. The risk is that the volume becomes the narrative: leadership sees a chart of thousands of blocked attempts and concludes the programme is working, while the handful of intrusions that matter arrive quietly through a valid account in a week when nothing was in the news. Report the noise separately from the credential-based and targeted activity, and never let the two appear on the same axis.
The objection worth taking seriously
The strongest objection is that "stalemate" is a defeatist framing that gives people permission to do nothing. The triage improvements are not trivial — teams genuinely investigate more in less time, junior analysts genuinely operate above their experience level, and small organisations genuinely get capability they could not previously staff. Calling that a stalemate understates a real gain and hands ammunition to every finance director looking for a reason to decline the renewal. That is fair, and the gain should be claimed. Any team that has cut investigation write-up time substantially should say so plainly. The distinction worth holding is between a productivity gain and a risk reduction. They are not the same thing and they are being reported as though they were. Faster, better-documented investigation of a longer queue of less serious items is a genuine improvement in the working life of a security team and an uncertain improvement in the probability of a material breach. The failure mode is not scepticism about the tooling; it is allowing an operational efficiency to be presented as a control, and then declining the unglamorous authentication and inventory work because the budget already went to the thing with the better demonstration. Take the triage gain, measure it honestly as a productivity number, and keep the risk conversation anchored to how intrusions actually start.
Common Questions
Are attackers using AI to write novel malware?
There is little credible evidence of new capability, and considerable evidence of faster production of existing tooling. Treat claims of AI-generated novel malware families with scepticism until the sample is published.
Should we enable autonomous response?
For reversible actions on well-understood detections, yes. For anything that disables an account or blocks a production service, keep a person in the path until you have reopen-rate data.
Does this change our security awareness programme?
Substantially. Anything teaching staff to spot poor language or obvious errors is now teaching an obsolete signal. Move the content to process anomalies and verification behaviour.
What should we expect over the next twelve months?
Expect security products to keep adding assistants faster than they add evidence, and expect procurement to start asking for evaluation data rather than accepting demonstrations. Expect the first serious incidents caused by over-trusted automated response to become public. Expect attackers to focus on voice and video rather than text, because text is now solved for them and the verification habits around a phone call are much weaker. And expect the honest defensive wins of the year to come from identity work that nobody will write a press release about.
AI Security Strategy Assessment — we separate the triage gains you can bank from the detection claims you cannot, and put the budget where intrusions actually start.
