Security teams spent the last two years reviewing agent pilots. The pilots ran in a sandbox, touched read-only data and produced a report that said the controls were adequate for the scope. That assessment is now being carried, unchanged, into production systems where the same agents write to ledgers, send messages to customers and call external services on a schedule nobody watches. The interesting risks in production agents are not the ones the pilot review looked for. They are the ones that only exist when something runs continuously, unattended, with credentials of its own.
A pilot fails by producing a bad answer. A production agent fails by producing a plausible one, at four in the morning, three hundred times
Here is what changes when an agent goes live, and what to control.
The four production-only risks
Standing authority. A pilot borrows a person's session. A production agent holds a credential permanently, which means its permissions are always live and its blast radius is whatever it can reach at three in the morning on a public holiday. Untrusted input reaching a decision path. The agent reads a supplier email, a customer ticket, a document, a web page. Any of those can contain instructions, and the agent has no reliable way to distinguish content it should act on from content it should merely read. Chained actions without a stopping condition. Agents call tools that call tools. In a pilot the chain is short and observed. In production it is long, conditional and capable of running until something external fails. Silent drift. Model updates, prompt changes, new data shapes and altered upstream formats change behaviour without a deployment event. There is no change control on a system whose behaviour depends on a provider's release schedule.
What actually works as a control
Rate and value limits per agent, per period — crude, unglamorous, and the single most effective containment available. Separate credentials per agent so an incident can be attributed and revoked without disrupting everything else. A hard boundary between reading untrusted content and taking consequential action, with the second requiring something the content cannot influence. And behavioural monitoring that alerts on volume and pattern rather than on individual actions, because the individual actions all look fine.
The control that gets skipped
The kill switch. Almost every production agent deployment lacks a tested way to stop it immediately without taking down the surrounding system, and the moment you need one is not the moment to build it. Test it quarterly, the way you would test a failover.
| Control boundary | Question to test |
|---|---|
| Identity and scope | Can this agent be attributed and revoked independently? |
| Action authority | Can external content influence approval for a consequential action? |
| Limits and stop | Are volume/value bounds enforced and shutdown rehearsed? |
| Change and monitoring | Can the team trace input, actor, outcome and behaviour changes? |
Qualitative summary of this article's source text, not a measured outcome or performance estimate.
Practical Guidance for Agentic Security Review
- Give every agent its own credential with its own scope.
- Set value and volume limits per agent, per period.
- Separate reading untrusted content from taking action.
- Build and test a kill switch before go-live.
- Monitor patterns and volumes, not individual actions.
- Treat model and prompt changes as changes requiring review.
- Log actor, input and outcome for every consequential action.
- Re-review anything promoted from pilot; the pilot assessment does not transfer.
The Regional Angle
The first regional consideration is that a large proportion of agent deployments in the Gulf are built and operated by systems integrators rather than internal teams, which means the security design is a procurement outcome. Credential scoping, rate limits and kill switches cost integrator time, are invisible in a demonstration, and will not appear unless they are written into the statement of work as acceptance criteria. The most effective security intervention available to a regional buyer is therefore a contract clause, which is not how security teams are used to working. The second concerns language, which creates an input-validation gap specific to bilingual operations. Agents here routinely process Arabic and English content, and the safety behaviour of most models is weaker outside English — instructions embedded in Arabic text in a supplier document are less likely to be recognised as an injection attempt than the same instruction in English. Test your agents against adversarial content in every language they will actually encounter, because the vendor almost certainly did not. The third is about the regulatory environment, which is moving faster than most agent deployments are documented. Financial regulators across the Gulf have been progressively tightening expectations on operational resilience, outsourcing and technology risk, and an autonomous process acting on customer data with a credential nobody can account for is exactly the kind of finding that generates a supervisory conversation. Regulated entities should assume they will be asked who authorised the agent, what it can do and how it is stopped — and should be able to answer in writing.
The objection worth taking seriously
The strongest objection is that this treats agents as a new category when they are just software, and that the discipline already exists. Service accounts with scoped credentials, rate limiting, input validation, change control, monitoring and emergency shutdown are all standard practice for any automated integration, and organisations have run unattended batch processes with production write access for decades without inventing a new security paradigm. Building a separate agent security programme risks producing a parallel set of controls that duplicates the one you have and is maintained less well. That is largely right, and the emphasis on reusing existing controls is correct. The genuine discontinuity is narrow but real: a batch process does exactly what it was coded to do, and its behaviour is a function of its version. An agent's behaviour is a function of its instructions, its inputs and a model that changes underneath it, which breaks the assumption every existing change control depends on — that behaviour changes only when someone deploys something. That single difference is what makes untrusted input dangerous and silent drift possible. So the answer is not a parallel programme; it is your existing controls applied with two additions: treat inputs as potentially adversarial in a way batch inputs never were, and treat provider-side changes as change events even though nobody on your side did anything.
Common Questions
Should an agent ever have write access to production systems?
Yes, within scoped operations, value limits and a tested stop mechanism. The alternative — a human executing everything the agent decides — removes most of the benefit and none of the risk of bad reasoning.
How do we handle prompt injection practically?
By architecture rather than by detection. Content that comes from outside should not be able to authorise an action; put the authorisation on a path the content cannot reach.
What is the most common gap?
Shared credentials, followed closely by the absent kill switch. Both are cheap to fix before deployment and expensive afterwards.
What should we expect over the next twelve months?
Expect the first significant incidents to involve standing credentials rather than clever attacks. Expect regulators in this region to begin asking about agent authorisation explicitly. Expect vendors to ship native rate limiting late. And expect pilot-era security assessments to be found attached to production systems during audits.
Agentic Security Review — we review what changed when your agents left the sandbox, starting with the credentials nobody can account for.
