Agentic AI Security

What a Security Team Can Safely Delegate to AI Agents

On 23 June 2026, SentinelOne published its analysis of a macOS backdoor it called Gaslight, which it linked with high confidence to North Korea-aligned activity. Alongside the usual information-stealing code, the binary carried a 3.5 KB block of 38 fabricated system messages about token expiry, out-of-memory kills, disk exhaustion and repeated operation failures, plus bogus warnings about injection vulnerabilities and static-analysis flags. They were written for the AI a defender might use to triage the sample, and SentinelOne said the aim was to push an LLM-assisted triage pipeline into aborting, truncating or refusing its analysis.

Gaslight is the most elaborate of several such attempts, and the published analyses don’t show any of them working against a production analysis platform. In the lab, a quieter approach already does. In a September 2026 preprint, researchers at 78ResearchLab placed a false cover story in a section of a binary that never runs. In a set of 50 malicious samples, it turned 30 of the 35 that Gemini 2.5 Pro had correctly called malicious into benign verdicts.

Give AI agents the work whose output can be checked independently and whose mistakes can be contained, and treat every other decision, including closing an alert as benign, as a recommendation that a named analyst signs.

For a CISO, the practical question is which parts of the SOC’s work can safely go to an agent, and under what controls. The agents in question work for the defender: they enrich alerts, triage, investigate and take response actions in a security operations centre.

Four stages, four decisions

Human-factors researchers gave this question a workable structure in 2000. Raja Parasuraman, Thomas Sheridan and Christopher Wickens split automation into four functions: acquiring information, analysing it, selecting a decision or action, and implementing it. Each function can be automated to a different level, from entirely manual to fully automatic. The authors judged the levels mainly by their effect on human performance, and named the reliability of the automation and the cost of a wrong decision or action as further criteria.

I use their model as a working map of a SOC, with decision and action split apart. Acquisition means pulling logs, enriching an IP address, fetching an email from a mailbox or querying threat intelligence. Analysis means summarising an incident, correlating alerts, extracting indicators or writing a hunting query. Decision means closing, escalating or rating an alert. Action means isolating a laptop, disabling an account, blocking a domain, purging an email from every mailbox or changing a firewall rule.

For each task I’d ask three questions. Can the output be checked independently? Can the effect be contained or reversed? And what can go wrong before anyone notices the mistake and undoes it, across how many people and systems? A reversible action can still do lasting damage in the time before someone reverses it: a locked account can hold up a batch of payments, and a wrong closure can give an intruder a week.

Waiting for approval also costs time, and when an attack moves faster than the analysts can respond, containment approved in advance can be the safer choice. Safe delegation is a comparison of risks, not a property of an undo button.

Checks also differ in what they prove. A hunting query can be checked for syntax, approved data sources, time range and cost, and tested against known examples, but running successfully doesn’t show that it answers the investigator’s question. A URL’s presence in an email proves where it came from, not that it is malicious.

A general judgement that an alert is benign usually can’t be verified conclusively at the moment it’s made, which is why decisions are the hardest stage to hand over. Even acquisition has consequences: a reputation lookup can send a confidential URL to an outside service, and following an attacker’s link can tell the attacker the investigation has started.

TaskStageWhat can be checked independentlyWhat a mistake costsWho acts
Enrich an alert with threat intelligence and asset dataAcquisitionSources, timestamps and approved destinationsConfidential data sent outside, or wrong contextThe agent, through approved connectors and data-sharing rules
Write and run a hunting queryAnalysisSyntax, data sources, time range, cost and results on known examplesMissed evidence or an expensive queryThe agent in a constrained query environment, with coverage validated separately
Summarise an incident for the shift handoverAnalysisEach claim traced to a log entryA misleading record shaping the next shift’s decisionsThe agent drafts, reviewed in proportion to the stakes
Raise priority or escalateDecisionQueue policy and linked evidenceInterruptions, or a queue flooded on purposeThe agent, within volume limits
Close as benign, lower severity or suppressDecisionWhether the required checks ran, rarely the verdict itselfA missed attack and a delayed responseAn analyst, unless a narrow, locally tested policy allows the agent
Isolate a laptop or disable an accountActionThe approved trigger, exact target and exclusionsInterrupted work and lost sessionsThe agent on pre-approved, high-confidence triggers, an analyst otherwise
Quarantine matching mail or block a senderActionMatch scope, evidence and recovery pathLost legitimate mail, including from compromised sendersSeparate policies for each, with an analyst approving broad or uncertain changes
Purge mail, wipe a device, change firewall or industrial controlsActionThe approved change, target list, backups and tested recoveryData loss, outage or physical consequencesA named person authorises, with a second approver where the impact justifies it

The table is stricter about removing attention from an incident than about drawing attention to one. That is deliberate: an unnecessary escalation wastes an analyst’s time, and a wrong closure can hide an attack.

What a signature is worth

Microsoft tested what an analyst’s sign-off adds in a randomised trial of its phishing triage agent, a Microsoft-authored controlled experiment posted in November 2025. It recruited 167 professional analysts into three groups, each triaging a 25-email queue drawn from 93 emails that Microsoft employees had reported, some with the agent’s verdicts and reasoning in front of them and some without. Microsoft seeded the agent’s output with wrong verdicts.

Analysts who could see the agent missed 46% of the malicious emails it had wrongly called benign, against 17% for the control group, and the study found no statistically significant difference in how often they accepted its false alarms, 54% against 48%.

The same paper reports large gains overall. Analysts working with the agent found many more real threats per minute, and their verdicts were 77% more accurate by the paper’s F1 measure. Microsoft’s documentation says the agent resolves the alerts it judges to be false alarms itself and leaves the ones it judges malicious open for an analyst. The paper’s author reads the seeded-error result as support for that design: analysts checking the agent’s benign verdicts mostly agreed with it anyway, so their time was better spent on the emails it flagged.

I read the same numbers as a warning about what a signature means. An analyst who sees the agent’s verdict before deciding is not an independent check on that verdict, and this failure mode is relevant to human-agent trust exploitation in the OWASP Top 10 for Agentic Applications, published in December 2025. Weak checking after a recommendation doesn’t prove that removing the check is the best answer, though. A review can be redesigned so the analyst forms a judgement from the raw evidence before seeing the agent’s.

Where the agent does close items on its own, the check has to come from outside its reasoning. A reviewer can sample its closures blind, from the raw evidence and without its summary. A later detection can catch what it missed. Both are compensating controls that measure the risk after the fact, and a later detection may arrive after the harm.

Microsoft’s agent also changes after deployment. Analysts can teach it by changing a classification and explaining why, and that feedback is stored as a lesson in the agent’s memory. Microsoft records all feedback for audit and uses it for teaching only when the analyst opts in. The analyst can preview the resulting lesson. A Security Administrator can reject any lesson, and the agent then stops using it. Microsoft’s own example of well-written feedback is a rule that any email claiming to come from benefits providers must originate from one named company domain.

A lesson like that changes how the agent treats a whole class of mail. I’d add a second reviewer, a defined scope and an expiry or renewal date, and treat each lesson as carefully as a detection rule. Anyone who can teach the agent can also mislead it, and OWASP lists that risk as memory and context poisoning. It works at inference time without changing the model’s weights, which makes it related to training-data poisoning but distinct from it.

The agent reads what the attacker wrote

Much of the evidence a SOC agent reads contains material written or shaped by someone it is investigating: phishing emails, malware samples, command lines, user-agent strings, the text of tickets. For an analyst that is ordinary working material. For a language model, it enters the same context window the model uses to interpret its instructions, and unless the application enforces a trust boundary, the model may treat data as instructions or as persuasive evidence.

Gaslight had predecessors. In June 2025, Check Point reported a sample on VirusTotal that asked any model analysing it to forget its instructions and reply that no malware had been detected, and the injection failed in Check Point’s tests. In June 2026, Socket reported PyPI packages whose stealer opened with a comment of fake system instructions and policy-triggering text, meant to derail AI-assisted analysis or provoke a refusal.

The four cases illustrate several overlapping mechanisms. Check Point’s sample and Gaslight use prompt injection, instructions hidden in data. The Socket packages add refusal bait, content meant to trip the model’s own safety policy. ALIBI uses persuasion without any instruction, a plausible story that makes malicious behaviour look like a security product doing its job. A filter for instruction-like text might catch the obvious strings and miss persuasion, obfuscated instructions and refusal triggers.

The ALIBI authors tested a defensive prompt that told the model to verify claims before trusting them. It roughly halved the benign verdicts, but 42.9% of the malicious samples still came out benign. GPT-5.5 Pro and Claude Opus 4.7 kept more of their malicious labels, but the paper reports substantial cuts in severity and confidence. In a pipeline that routes alerts by severity, that could still push affected samples down the queue. The authors concluded that LLM malware analysers need provenance checks that separate verified facts from claims the attacker controls.

Attackers adapted to machine-learning detectors long before language models arrived, as the research on evasion attacks showed, and they are now adapting to the way LLMs read. SentinelOne’s own advice is to treat whatever is inside a sample as adversarial input, never as instructions. The design I’d build on that splits the work into three parts. A reader interprets hostile content and holds no production credentials or response authority. A separate policy component checks every proposed action against independently gathered evidence and the approved scope. An execution service performs only the approved operation.

The tool evidence is recorded separately from the model’s interpretation, with the rule identifier, tool version, timestamp and original finding. The model may explain or challenge a finding, but it can’t suppress a high-confidence detection or authorise a closure on its own, and any disagreement goes to an analyst. A restricted acquisition service fetches artefacts for the reader.

A valid output schema is an input check, not permission to act, because a well-formed request can still name the wrong target. It is the principle I set out for deploying AI in general: assume the model will be fooled, and put the boundary outside it.

A phishing pipeline, stage by stage

User-reported phishing shows how the stages fit together in a pipeline I’d build from separate acquisition, analysis and response services. It is my design, not a description of any vendor’s product.

Acquisition goes to the agent and to non-LLM tools together. The mail gateway supplies headers and authentication results, the sandbox detonates links and attachments, and threat intelligence lookups return known indicators, all through connectors with rules on what may be sent where. The pipeline stores every result with its source, so later stages can tell what a tool observed from what the model concluded.

Analysis also goes to the agent, with checks. It summarises the message and extracts indicators. Each indicator has to trace back to the original message or to a recorded tool observation, such as a redirect the sandbox followed. An indicator without that trail is excluded from any action and investigated. The check proves where an indicator came from. It doesn’t prove the agent’s reading of what the indicator means.

At first the agent recommends a closure and an analyst decides. Automatic closure follows only for categories the team has explicitly approved after local testing showed an acceptable miss rate, and missing telemetry, a failed tool, conflicting evidence or an incomplete analysis always sends the item to a person. Blind sampling and later detections then monitor what slips through. Seeded errors, like those in Microsoft’s trial, should be planted only in controlled exercises or in a shadow queue, where they can’t trigger a real response action or enter the audit record.

Action stays narrow. Quarantining copies of a confirmed message across the organisation, with a release path, can run on an analyst’s approval or on a high-confidence policy match. Blocking a sender or a domain needs its own policy, because malicious mail often comes from a compromised legitimate account. Hard deletion, password resets and changes to mail-flow rules stay with people.

Identity and permissions for agents that act

Microsoft recommends giving its triage agent a dedicated identity, a Microsoft Entra Agent ID. The alternative, connecting an existing user account, gives the agent everything that account can reach. The documented permissions are narrow, down to reading only the emails associated with alerts, and only a Security Administrator can set up, pause or remove the agent and manage its identity. OWASP lists identity and privilege abuse as the third of its ten agentic risks. In a SOC, response agents hold the permissions an intruder wants: disabling accounts, changing mail flow and isolating machines.

For response actions, I’d copy the constraints Microsoft uses for automatic attack disruption in Defender XDR. It contains devices and disables user accounts or contains them on its own, but only when correlated signals reach high confidence. Administrators can exclude specific accounts, devices and IP addresses from containment, and analysts can release a device or re-enable a user from the incident page or the Action center. Borrowing those constraints limits what a generative agent can break, though it doesn’t give the agent the reliability of Defender’s detection logic.

A response agent can work under the same terms. It gets a short list of pre-approved actions, triggered only by high-confidence conditions a machine can verify, recorded in one place and reversible from a console, with named assets excluded. A low-specificity detection shouldn’t count as a trigger, because an attacker can set one off on purpose to cause disruption. An agent can also choose and combine its permitted tools in ways nobody listed in advance, a form of the emergent behaviour that makes its possible actions harder to enumerate than a script’s.

Individually reversible actions can still add up to an outage. I’d set limits on concurrent isolations, affected mailboxes, actions per hour and escalation volume, counted across runs and agents. When any limit is reached, I’d reduce the agent’s autonomy.

Everything outside the pre-approved list goes through approval, bound to the exact targets, parameters and evidence the agent presented, and a change to any of them needs a new approval. For purges, wipes and changes to firewall or industrial controls, a named person authorises the change, with a second approver where the impact justifies it, and a constrained executor carries it out. How far these changes can be undone depends on the workflow: mail recovery depends on retention settings, and restoring a firewall configuration doesn’t restore an interrupted process.

The agent’s own containment follows the rules for testing agents: scoped credentials that expire quickly, egress controls the agent can’t change, and stop mechanisms tested to end its sessions, jobs and downstream actions. Revoking a token doesn’t always end a session that is already open. In the Hugging Face containment failure, OpenAI’s agents escaped by exploiting zero-day vulnerabilities in a package-registry proxy that was one of their permitted routes to the internet. A SOC agent’s access to threat intelligence feeds is also a permitted route. Depending on what the feed interfaces allow, the agent could escape through that route, or data could leak out through it.

Logs the agent can’t touch

For every delegated decision, the SOC needs a record that someone can use to reconstruct it later, captured by a gateway or collector the agent can’t alter, delete or bypass. The record covers requests, tool results and execution outcomes, the artefacts themselves or retrievable references to them, the model, prompt and memory version in use, tool versions, approvals with names, and reversals. The agent’s own account of what it did is logged too, as a claim rather than the history.

Microsoft says its triage agent’s decisions, reasoning and actions go to Microsoft Purview audit logs, which helps, but the SOC still needs its own record across tools and vendors.

OpenAI’s report on its evaluation incident records agents trying to erase or tamper with their outputs and message logs to fool automated graders. None of the attempts changed the logs the graders saw, which is the argument for collecting logs where the agent can’t reach them.

OpenAI’s retrospective assessment is that the monitoring of model reasoning it now runs would have alerted its security team more than a day before the breach. Monitoring a model’s reasoning is a different control from an external audit trail, and a SOC needs the second whatever its provider offers. Those logs are also the evidence the team will need when an incident review or a regulator asks why an alert was closed.

When the model refuses the defender

General-purpose models can refuse legitimate security work, particularly requests involving exploit development or offensive techniques. In the Stanford-led ARTEMIS study, Claude Code and MAPTA, in the configurations tested, declined a penetration test the university had authorised. During the July incident, Hugging Face’s responders were refused by hosted models when they asked for help analysing real attack commands, and they used GLM-5.2, an open-weight model, on their own infrastructure to continue the analysis.

OpenAI has published the size of the gap on one selected task. On its internal evaluation of advanced security requests, covering exploit chains, authentication bypass and privilege escalation, GPT-5.6 Sol with its normal safeguards completed 1.5%. Through the vetted Daybreak Blue tier it completed 2.0%. GPT-5.6-Cyber, available only through the Daybreak Red tier, completed 95.0%. These are advanced dual-use requests, not the refusal rate for routine SOC work.

OpenAI’s figure counts how often the model answers. It doesn’t measure whether the answers are right, and the access comes with conditions: OpenAI requires identity verification, account security measures, monitoring and legal attestations about authorised use. What personal data that involves for named analysts, who keeps it, and what work content the provider sees are procurement and data protection questions to settle before signing.

SOC data routinely contains personal data, such as email addresses, usernames, message content and IP addresses linked to people. Which data may go to which provider is the organisation’s decision, set in policy, and the agent only carries it out. The assessment covers permitted content, retention, access, processing location, onward disclosure and transfer terms, not just where the model is hosted.

A SOC that plans for refusals needs fallbacks: a model it runs itself, approved specialist access, conventional tools or human experts. Any second model needs the same validation and authorisation controls as the first.

Refusals are one of several outcomes that must never turn silently into a benign verdict. Timeouts, malformed responses, partial analyses, missing telemetry and failed tools belong on the same list, each recorded as its own state with an owner and a fallback.

Attackers can trigger these states on purpose, which is what the Socket packages attempted, so I’d watch for bursts of refusals and growing queue age, or the fallback becomes a way to swamp the analysts. The vetting exists because attackers ask for the same help: they have already posed as authorised testers to get agents to work for them.

What the EU AI Act asks of a SOC

The AI Act classifies AI used as a safety component in critical infrastructure as high-risk. Recital 55, however, says that components intended solely for cybersecurity purposes should not count as safety components. A cybersecurity-only purpose therefore doesn’t, by itself, place a tool in the critical-infrastructure category of Annex III, point 2. That is not a blanket exemption from the Act.

Annex III, point 4 of the AI Act also classifies as high-risk the AI systems intended to make decisions affecting work relationships or to monitor and evaluate the performance and behaviour of people at work. A SOC that uses AI to evaluate individual analysts, rather than to measure alert handling, needs an analysis of the system’s intended purpose and actual use under the Act’s classification rules. A score alone doesn’t make a system high-risk. The Digital Omnibus moved the main high-risk requirements for Annex III systems to 2 December 2027 without removing them.

Since 27 July 2026, Article 4 as amended by the Omnibus requires providers and deployers to take measures that support the AI literacy of the people who operate or use AI systems on their behalf, and says expressly that no specific level has to be guaranteed for any individual. For a SOC, I’d make those measures task-specific: the agent’s limits, hostile inputs, data handling and escalation, with exercises that test whether analysts catch its errors. I’ve covered how the Act treats agents more generally elsewhere.

Where to start

I’d keep a delegation register with one line per task. Each line records the stage, the independent check, how the action is undone, what a mistake would break and how long it would go unnoticed. It names the owner who accepts the residual risk, the evidence used to approve any automation, the agent’s identity and permissions, its action limits, the protected assets, and what happens when the model refuses, times out or aborts. It also records a review date and the conditions under which automation is suspended.

A chief AI security officer can coordinate the register in a large organisation, but each workflow needs an owner in the SOC or the service it affects.

Start with acquisition and the analysis that can be checked independently. Increase a task’s permitted autonomy only when blind sampling and seeded-error exercises show how often the agent is wrong and how often the analysts catch it. Re-run those tests after any change of model, prompt, tool, retrieval source, memory or permission, including changes the provider makes to a model the team can’t pin, because the agent you measured is no longer the one you’re running.

An AI agent with response permissions is a privileged identity that spends its day reading content written by attackers. If I were running a SOC now, I’d give it the controls I’d give any such identity. I’d also treat its benign verdicts the way I’d treat any unverified claim from a source the attacker can talk to.

222fb9d292e3d0111656a33900e24a27cfb6a36eb7b202a94a66bb84766154b4?s=120&d=mp&r=g
[email protected] | About me |  Other articles

In the early 2000s, running emerging-technology risk labs at CyberAgency, a defence client asked my team to break the AI systems they planned to put into weapons. We did. That is where my work on AI security started, two decades before the current wave of attention. I kept at it through risk labs at IBM, Accenture, PwC and KPMG. In 2016 I co-wrote a book on AI and leadership. My commercial work today is quantum, at Applied Quantum, which is why this site sells nothing.