Cybersecurity / Source date:

AI-Generated Phishing Removes the Obvious Tells

Fluent, contextual lures defeated the grammar-based detection habits users were trained on.

Illustration of a hand lifting a telephone receiver beside a closed verified-contacts directory.

In March, Europol published an assessment of how criminals are using large language models. This spring's internet crime figures put business email compromise losses at roughly 2.7 billion dollars for last year, against total reported losses of 10.3 billion — an order of magnitude above ransomware, and almost entirely unglamorous. Neither document is the reason this is worth writing about now. The reason is what has changed in the mailbox over the past quarter. The messages are fluent. They are correctly formatted. They reference a real project, use your internal vocabulary, and arrive at a plausible moment in a real conversation.

The phishing email you should worry about is not the one that arrives from nowhere. It is the reply to a message you actually sent

Thread hijacking is the technique that matters, and it predates the current wave of generative tools. An attacker with access to one mailbox — a supplier's, a partner's, a colleague's — replies inside an existing conversation. The subject line is one you wrote. The quoted history below is genuine. The request is new. What has changed is the writing. Previously the reply arrived in the wrong register: too formal, oddly phrased, plainly not the person whose name was on it. That mismatch was the tell, and it was the only reliable one most staff had. It has now been removed. A model given the thread above can match tone, length, greeting habits and the particular brusqueness of the person being impersonated.

Four tells that have stopped working

Language quality. Spelling and grammar errors are gone. The long-standing theory that poor writing was deliberate filtering for gullible targets was always partly true and is now irrelevant, because the cost of good writing has fallen to zero. Formatting and branding. Correct logos, correct footers, correct disclaimer text, correct font. Reproducing a corporate template is trivial. The generic salutation. Messages now use your name, your role, your project and your colleague's name, assembled from material that is publicly available or already in the compromised thread. Urgency mismatch. The classic version was implausible panic. The current version has a reason — a stated approval deadline, a payment run, a closing — and is willing to wait three days for the right moment. Every one of those four is what your awareness training teaches people to look for.

What still works as a signal

Move detection off the text and onto the metadata and the request. A change of channel. A request to continue on a personal address or a messaging app remains one of the strongest indicators available. A new payment instruction. Any first-time account detail, any change to existing details, any unfamiliar beneficiary. An unusual authority path. Someone senior approaching someone junior directly, bypassing the person who would normally handle it. Identity infrastructure. The sending domain, its registration age, authentication results, whether the display name matches the actual address. None of this is improved by better prose. A request that sits outside the recipient's normal process. The attack almost always requires someone to do something they do not usually do. That list has one property in common: none of it requires the reader to judge whether the writing sounds right.

The awareness programme has a measurement problem

Simulation click rates were always a weak metric. They measured how obvious your simulations were, and they punished the people who click links for a living. They are now actively misleading. A programme reporting a two per cent click rate against templates containing the old tells is measuring nothing, because the real messages no longer contain them. The number will look excellent right up until the incident. The instruction has to change as well. "Spot the fake" is no longer a skill you can train, because the observable differences have been removed. "Run the verification step whenever the request meets these conditions, regardless of how legitimate the message looks" is trainable, testable and does not depend on anyone's judgement about tone. Measure report rate and verification compliance. Retire click rate as a headline number, and explain to the board why it went up.

Three controls worth more than any training

Out-of-band verification bound to a stored contact record. The verification must use contact details held in your own system, never details supplied in the message. An attacker who controls the thread also controls the phone number in the signature block. Payment detail changes as a dual-control process. Any change to supplier bank details requires callback to a known number, a second approver, and a written record. This one control removes most of the loss exposure and costs nothing but discipline. Technical hygiene that surfaces identity. Strict domain authentication and alignment, external-sender marking prominent enough to be noticed on a phone, display-name lookalike detection, and alerting on mailbox auto-forward rules — which is how an attacker keeps reading the thread after the password change.

Verify the request outside the messageA defensive workflow summarised from this article. It is not a fraud-loss estimate or guarantee of prevention.
  1. Recognise a process change

    Treat new payment instructions, channel changes and unusual authority paths as verification triggers.

  2. Use a stored contact

    Retrieve the callback number from your own records, not the email or signature.

  3. Apply dual control

    Have a second approver review changes to supplier bank details.

  4. Keep the evidence

    Record the verification and authorisation before the change is used.

Qualitative summary of this article's source text, not a measured outcome or performance estimate.

Voice is next, and it needs a rule now

The same capability applies to audio, and the sample requirement is short. A recorded conference presentation, a podcast appearance or a voicemail greeting is sufficient material. The rule to establish before it matters: a voice on a phone call is not authentication. Instructions received by voice are unverified until confirmed through the stored-contact process, exactly like an email. Write that down while it still sounds excessive.

Practical Guidance for Phishing Defense Modernization

  • Stop training people to spot bad grammar. It no longer exists.
  • Bind verification to stored contact records, never to details in the message.
  • Make bank detail changes dual control with callback to a known number.
  • Alert on mailbox auto-forward rules and unexpected inbox rule creation.
  • Enforce domain authentication and make external sender marking visible on mobile.
  • Measure report rate and verification compliance, not click rate.
  • Treat channel change requests as the primary trigger for suspicion.
  • Agree now that a familiar voice is not authentication.

The Regional Angle

Three factors make this land differently here, and the first is the dominant local pretext rather than the corporate one. Impersonation of government and authority is the region's most productive attack theme, well ahead of the chief-executive wire transfer that dominates the international literature. Visa and labour fine notices, customs clearance demands, traffic fine portals, tax authority correspondence and utility disconnection warnings all work here because the underlying processes are real, frequently involve genuine payments to genuine portals, and are administered by entities most staff cannot readily verify. Until recently these arrived in poorly translated Arabic or clumsy English, which gave recipients a chance. That has now closed: the same messages arrive in fluent, correctly registered formal Arabic. The rule to embed is absolute and simple — no government portal is ever reached from a link in a message. Finance teams and public relations officers keep a bookmark list of the real portals and use nothing else. Worth adding one uncomfortable observation: most awareness programmes in the region run English-only simulations, which means their reported click rates are measuring resistance to attacks in the language staff are least likely to fall for. The second is an intermediary that sits outside almost every supplier verification process. Public relations officers, typing centres and document-clearing agents handle a constant stream of government transactions on behalf of companies here, and payments to them are frequent, irregular in amount, poorly documented and often settled quickly because a visa or licence deadline is looming. That combination — legitimate urgency, variable amounts, weak paperwork, and a counterparty who is not in the vendor master — is an almost perfect target profile, and these payments are usually excluded from whatever supplier verification discipline the finance team has built. Bring them inside it. Give the clearing agents proper vendor records with verified bank details, apply the same callback rule to any change, and treat an unexpected payment request from a document-clearing intermediary with the same scepticism as one from an unfamiliar supplier. The third is the approval culture in owner-led and family groups, where the attack works because the described behaviour is genuine. In many regional businesses a direct message from the principal really does move money, quickly, without a documented approval chain, and asking for confirmation can be read as questioning authority. That makes message and voice impersonation unusually effective here, and it is not a problem the finance team can solve on its own. The callback-to-stored-number rule only operates if the principal endorses it publicly, in front of the people who will have to apply it, and ideally tests it once. Without that endorsement the rule exists on paper and is bypassed the first time it is needed.

The objection worth taking seriously

The strongest objection is that this is a training-and-process argument dressed up as news. Filters already intercept the overwhelming majority of these messages before anyone sees them. Sophisticated business email compromise was always well written — the people stealing millions were not sending messages full of typos — so the marginal gain to a competent attacker from better language tooling is close to nothing. The controls recommended here are the controls that were recommended five years ago, and the only genuinely new element is the vocabulary. That is largely correct and worth conceding without hedging. The high end of this crime has not changed, and none of the advice above is novel. What has changed is the middle. The volume tier — the operators sending thousands of generic messages that used to be discarded on sight — can now produce output indistinguishable from the high end, at the same cost per message as before. The consequence is not a new kind of attack but a change in the distribution: a much larger share of what reaches an inbox is now good enough to work. And that breaks the control everyone has quietly relied on, which was human judgement applied message by message. That judgement was always the weakest layer; it is now unreliable at scale. Which is precisely why the answer is process controls that do not require anyone to assess a message at all.

Common Questions

Does this mean awareness training is pointless?

No, but its content has to change. Stop teaching detection cues that no longer exist and teach the verification behaviours that do not depend on judging a message.

Will email filtering catch these?

Most of them, most of the time. Filtering struggles specifically with thread hijacking from a genuinely compromised legitimate account, which is the technique this article is about.

What is the single highest-value control?

Callback verification on supplier bank detail changes, using a number held in your own records. It addresses the largest category of financial loss and requires no technology purchase.

What should we expect over the next twelve months?

Expect purpose-built criminal tooling to be packaged and sold openly within months, because the demand is obvious and the underlying capability is widely available. Expect voice cloning to appear in finance fraud cases before the end of the year, initially as a confirmation step on an email request rather than as the primary channel. Expect insurers to start asking specific questions about payment verification controls at renewal, and to price accordingly. And expect detection vendors to pivot their messaging from content analysis to behavioural and relationship signals, because content analysis is the part that has stopped discriminating.


Phishing Defense Modernization — we replace the detection cues that stopped working with verification steps that do not depend on judgement, and we start with the payment process that carries the loss.

Continue reading

Talk to OPS

Start with the operating problem.