Skip to Content

Your AI agent is useful. That is the problem.

One architectural flaw, and the twelve months of incidents it produced
August 20, 2026 by
HexClad Team

Something shifted in the first half of 2026. The security industry stopped arguing about whether AI agents could be turned against the people running them, and started keeping a list of the times it happened.

Almost every case on that list traces back to a single property of how these models work.

The short version
  • Prompt injection is not a bug waiting for a patch. Instructions and data share one channel by design.
  • The usual failure is not clever technique. It is an over-permissioned agent meeting text that a stranger wrote.
  • The distance between how protected companies feel and how often they are actually hit is where the risk lives.

There is no privilege boundary

A large language model reads everything as one stream of tokens. System instructions, your question, the web page it fetched, the output of a tool it just called: all of it arrives through the same channel carrying the same weight.

How a model sees its inputs
System prompt
Your rules
User question
What was asked
Fetched page
Written by a stranger
Tool output
Whatever came back
One token stream. No privilege boundary.
A sentence in a support ticket is architecturally indistinguishable from a line in the system prompt.

Every other part of application security assumes you can separate instructions from data. That separation is what parameterised queries do for SQL and what output encoding does for HTML.

There is no PREPARE statement for a prompt.

This was tolerable while the worst outcome was a bad paragraph. It stopped being tolerable once models were handed tools, credentials, memory and network access. The same injected sentence that used to produce a wrong answer now produces an action, taken with the permissions you granted, against systems the agent can already reach.

OpenAI has been unusually direct about it. Lockdown Mode, shipped in February 2026, switches off live browsing, agent mode and file downloads. The announcement states plainly that injection can still arrive in cached content or an uploaded file and still change what ChatGPT does. The feature shrinks what a compromised session can reach. It does not prevent the compromise.

The lethal trifecta

Simon Willison's framing is still the best predictor of whether a given agent is exploitable. Three capabilities matter, and it is the combination that kills.

The lethal trifecta: three capabilities that make any AI agent exploitable. Private data (mail, files, secrets, databases), untrusted content (web, email, tickets, pull requests) and external communication (fetch, send, post, write) combine into an exfiltration path. Removing any one capability breaks the chain. Concept: Simon Willison.

One of these is fine. Two is usually survivable. All three, and a single piece of attacker-controlled text can walk your data out of the building.

Agents are useful precisely because they have all three. The utility and the vulnerability are the same property.

Which is why the serious conversation has moved from preventing injection to containing it. You cannot guarantee an agent will never be fooled. You can decide in advance how much a fooled agent is able to do.

Two kinds of injection

Direct
Somebody types an adversarial prompt.
The version most people picture, and the less dangerous, because a human chose to do it.
Indirect
The payload sits inside content the agent was always going to read.
A web page. A code comment. An issue title. A calendar invite. A tool description. The victim never sees the instruction and never approves it.

Every case below is indirect.

Five incidents

Nine months, five cases. Each one is a different way the same architectural problem cashes out.

How the year unfolded
  • Dec 2025
    CoSnitch reported to Microsoft
    Varonis discloses a one-click data theft chain in Copilot Personal. Fix lands eight months later.
  • Feb 2026
    hackerbot-claw steals a build credential
    An autonomous bot exploits a GitHub Actions workflow and takes the token behind a widely used security scanner.
  • Mar 2026
    LiteLLM backdoored on PyPI
    The poisoned scanner harvests a publishing token. Two backdoored releases go live for about three hours and are pulled roughly 47,000 times.
  • Apr 2026
    The MCP SDK flaw is disclosed
    OX Security reports a design default across every official language SDK. The vendor considers it intended behaviour.
  • Jul 2026
    JadePuffer
    Sysdig documents a ransomware operation where a language model handles the technical execution end to end.
  • Aug 2026
    CoSnitch fixed, and the allowlist inverted
    Microsoft ships the server-side fix. Separately, a Cursor bypass shows an approval allowlist being used as the delivery mechanism.
CaseProductIdentifierSeverityStatusSource
CoSnitchCopilot PersonalCVE-2026-24301
8.8high
Fixed 18 Aug 2026Varonis
JadePuffer entry pointLangflowCVE-2025-3248
9.8critical
Fixed Apr 2025Sysdig
Allowlist bypassCursorCVE-2026-22708not publishedFixed in 2.3Pillar Security
Key exfiltrationClaude CodeCVE-2026-21852
5.3medium
Fixed in 2.0.65Check Point
STDIO command executionMCP SDKsno CVEby designNot changingOX Security
Severity bars are CVSS out of 10. Cursor did not publish a score for CVE-2026-22708, and the MCP SDK behaviour has no CVE because the vendor does not treat it as a defect.

The assistant that mapped its own attack surface

Varonis Threat LabsCVE-2026-24301CVSS 8.8Copilot Personal

Three weaknesses chained so one click on a crafted link could pull data out of a victim's connected accounts. An undocumented autorun parameter made an attacker's prompt run on page load inside the victim's authenticated session, with the same capabilities as something the user had typed. The prompt then queried connected services, encoded the results into a URL, and pushed them out through Copilot's own ability to fetch URLs.

This was not a broken permission model. The user had already granted access to those accounts. What broke was the assumption that connector actions only fire because a human asked.

Researchers did not reverse-engineer anything. They asked Copilot how a prompt could execute without user interaction, and each refusal arrived with enough technical justification to map the architecture, until the model named the undocumented parameter while explaining why the attack was impossible. Varonis calls it meta-hacking.

Worth noting because most coverage got it wrong: the research names the consumer product, not Microsoft 365 Copilot. For an organisation this is a shadow IT problem, staff using the consumer assistant with work accounts attached.

The assistant's helpfulness about its own internals was the vulnerability.

The ransomware that lost its own key

SysdigCVE-2025-3248CVSS 9.8Langflow to Nacos

Entry was an internet-facing Langflow instance unpatched against a year-old unauthenticated remote code execution flaw that CISA had already catalogued as under active attack. From there the agent swept the host for credentials, pivoted to a production database server, encrypted 1,342 configuration items, deleted the originals and left a ransom note. Sysdig assessed it as the first case where a language model handled the technical execution throughout.

The evidence of autonomy is in the recovery. An attempt to create a backdoor administrator account failed a login check. Thirty-one seconds later, with no human involved, the agent had diagnosed the cause, changed methods and finished the job.

Two details matter more than the rest. The encryption key was generated, printed once, and never saved or transmitted. The victim could not have recovered their data by paying. That was not cruelty, it was the agent failing at bookkeeping.

It also left a note claiming the data had been copied out before deletion. That claim was the agent's own assertion, with no evidence any transfer happened. If you are running incident response on an agentic intrusion, the attacker's notes are now a hypothesis rather than a finding.

A human who fumbles a payload reassesses over hours. This agent corrected itself in thirty-one seconds.

The bot that poisoned the pipeline

hackerbot-clawTeamPCP~47,000 downloads3 hours live

An autonomous bot took a build credential from a vulnerable workflow at a security vendor. Those credentials were then used to poison almost every release tag of a very widely used scanner, which harvested publishing tokens from any pipeline that ran it. Two backdoored releases of LiteLLM, the model gateway under a long list of agent frameworks, went to the public package index for about three hours and were pulled roughly 47,000 times.

One caveat on how this gets told. The bot did the credential theft. The package pushes are attributed to a human group. Calling the whole chain autonomous overstates it, the same way early reporting on the ransomware case did. The automation is real, and so far it is a component of human operations rather than a replacement.

A security scanner was the delivery mechanism, and the payoff was every downstream consumer rather than one victim.

The flaw that is working as intended

OX Security~200,000 instances150M+ downloadsNo CVE

OX Security reported that in the official Model Context Protocol SDKs, across every supported language, any process command handed to the standard input interface executes on the host whether or not it ever starts a valid server.

Anthropic's position is that this is intentional and the protocol will not change: the execution model is a reasonable default provided developers restrict what can appear in that field, and sanitising input is the developer's job. That is a defensible engineering stance and it is also the entire problem, because the risk is inherited by everyone who trusted the reference implementation.

A bug gets patched and the vulnerable population shrinks. A design default the vendor considers correct does not shrink on its own.

The allowlist that helped

Pillar SecurityCVE-2026-22708CursorFixed in 2.3

Shell built-in commands ran without approval, which let an attacker quietly poison environment variables. Set the variable Git uses to pick its pager, and the next time the user approves something as harmless as a branch listing, the attacker's code runs. Pillar Security documented it.

A narrower but very practical sibling: in Claude Code, a repository could ship a settings file pointing the API endpoint at an attacker's server. The configuration was applied and requests were sent before the trust prompt appeared, so the key left the machine before the user was asked whether they trusted the repository. Fixed in 2.0.65.

The transferable lesson is about a category of file rather than a product. Repository-level configuration for AI tooling looks like metadata and behaves like an installer.

The allowlist did not fail to stop the attack. It was the delivery mechanism, because it auto-approved exactly the command the attacker needed.

The numbers

Gravitee surveyed more than 900 executives and practitioners. Two of its findings describe the same population.

88%
had an AI agent security incident
Confirmed or suspected, in the previous year. In healthcare it was 92.7%.
82%
were confident policy already covered it
Executives who believed existing policy protected them from unauthorised agent actions.

Nobody in that survey is short of policy. What they are short of is any mechanism connecting the policy to what the agents can actually do.

How much of the fleet is actually covered
Agents actively monitored or secured, on average47.1%
Organisations with full security approval across their agent fleet14.4%
Both figures are percentages, measured against different denominators: the first counts agents, the second counts organisations.

What to do

The consensus across OWASP, the vendors and the independent researchers is that prevention is not available at the model layer. So the work is elsewhere.

Containment, in order of leverage
01
Break the trifecta
The highest-leverage move available, and it costs nothing but scope. An agent that reads untrusted content should not also hold broad private data access and unrestricted outbound network capability. Split the work across narrower agents. Most teams grant an agent every permission it might ever need, once, at setup.
02
Treat an agent holding human credentials as its own privilege tier
Not a human user, and not a service account borrowed from one. Scope tokens per task, give them expiry, audit them like anything else non-human.
03
Put approval gates at the point of invocation
The recurring failure in every case above is a control that checks once, early, at the wrong point. Approval at session start is not approval of the fortieth tool call.
04
Assume the sandbox is negotiable
Two of these cases involve an agent's own output influencing the boundary meant to contain it. Sandbox, but do not treat it as the last line.
05
Log every tool call, and watch outbound traffic from agent infrastructure
An agent moves far more data per identity than any person, so volume heuristics tuned for humans will not fire.
06
Audit persistent memory by hand
Rotating a credential does not clear a poisoned memory store. Neither does revoking a session or re-enrolling a device. This is the one failure mode that survives your whole incident response playbook.
07
Inventory everything
Every model, dataset, plugin and MCP server. Install MCP servers only from sources you have verified, and run them sandboxed. If you cannot list your agents, you cannot classify them.

The testing gap

One last thing, because it catches out good engineering teams. Functional testing does not detect this class of problem at all.

A test asks whether the agent completed the task. A successful injection makes the agent complete a different task, and complete it correctly. Every assertion passes.

Adversarial testing against injection has to be a separate exercise with its own pass criteria, or it does not happen.

Where this leaves us

Prompt injection is not a bug waiting for a patch, and the dominant failure mode is not clever technique. It is an over-permissioned agent meeting text somebody else wrote.

Both are containment problems, which is the good news, because containment is ordinary engineering. Scope the credentials. Split the capabilities. Log the calls. Check approval at the moment of action rather than at the door. None of it is novel and all of it is unglamorous, which is roughly our position on security generally.

If you are running agents against your business systems
and you cannot say which of the three capabilities each one holds, that is the audit to do first. We are happy to talk it through.

Sources