Something shifted in the first half of 2026. The security industry stopped arguing about whether AI agents could be turned against the people running them, and started keeping a list of the times it happened.
Almost every case on that list traces back to a single property of how these models work.
- Prompt injection is not a bug waiting for a patch. Instructions and data share one channel by design.
- The usual failure is not clever technique. It is an over-permissioned agent meeting text that a stranger wrote.
- The distance between how protected companies feel and how often they are actually hit is where the risk lives.
There is no privilege boundary
A large language model reads everything as one stream of tokens. System instructions, your question, the web page it fetched, the output of a tool it just called: all of it arrives through the same channel carrying the same weight.
Every other part of application security assumes you can separate instructions from data. That separation is what parameterised queries do for SQL and what output encoding does for HTML.
There is no PREPARE statement for a prompt.
This was tolerable while the worst outcome was a bad paragraph. It stopped being tolerable once models were handed tools, credentials, memory and network access. The same injected sentence that used to produce a wrong answer now produces an action, taken with the permissions you granted, against systems the agent can already reach.
OpenAI has been unusually direct about it. Lockdown Mode, shipped in February 2026, switches off live browsing, agent mode and file downloads. The announcement states plainly that injection can still arrive in cached content or an uploaded file and still change what ChatGPT does. The feature shrinks what a compromised session can reach. It does not prevent the compromise.
The lethal trifecta
Simon Willison's framing is still the best predictor of whether a given agent is exploitable. Three capabilities matter, and it is the combination that kills.
One of these is fine. Two is usually survivable. All three, and a single piece of attacker-controlled text can walk your data out of the building.
Agents are useful precisely because they have all three. The utility and the vulnerability are the same property.
Which is why the serious conversation has moved from preventing injection to containing it. You cannot guarantee an agent will never be fooled. You can decide in advance how much a fooled agent is able to do.
Two kinds of injection
Every case below is indirect.
Five incidents
Nine months, five cases. Each one is a different way the same architectural problem cashes out.
- Dec 2025CoSnitch reported to MicrosoftVaronis discloses a one-click data theft chain in Copilot Personal. Fix lands eight months later.
- Feb 2026hackerbot-claw steals a build credentialAn autonomous bot exploits a GitHub Actions workflow and takes the token behind a widely used security scanner.
- Mar 2026LiteLLM backdoored on PyPIThe poisoned scanner harvests a publishing token. Two backdoored releases go live for about three hours and are pulled roughly 47,000 times.
- Apr 2026The MCP SDK flaw is disclosedOX Security reports a design default across every official language SDK. The vendor considers it intended behaviour.
- Jul 2026JadePufferSysdig documents a ransomware operation where a language model handles the technical execution end to end.
- Aug 2026CoSnitch fixed, and the allowlist invertedMicrosoft ships the server-side fix. Separately, a Cursor bypass shows an approval allowlist being used as the delivery mechanism.
| Case | Product | Identifier | Severity | Status | Source |
|---|---|---|---|---|---|
| CoSnitch | Copilot Personal | CVE-2026-24301 | 8.8high | Fixed 18 Aug 2026 | Varonis |
| JadePuffer entry point | Langflow | CVE-2025-3248 | 9.8critical | Fixed Apr 2025 | Sysdig |
| Allowlist bypass | Cursor | CVE-2026-22708 | not published | Fixed in 2.3 | Pillar Security |
| Key exfiltration | Claude Code | CVE-2026-21852 | 5.3medium | Fixed in 2.0.65 | Check Point |
| STDIO command execution | MCP SDKs | no CVE | by design | Not changing | OX Security |
The assistant that mapped its own attack surface
Three weaknesses chained so one click on a crafted link could pull data out of a victim's connected accounts. An undocumented autorun parameter made an attacker's prompt run on page load inside the victim's authenticated session, with the same capabilities as something the user had typed. The prompt then queried connected services, encoded the results into a URL, and pushed them out through Copilot's own ability to fetch URLs.
This was not a broken permission model. The user had already granted access to those accounts. What broke was the assumption that connector actions only fire because a human asked.
Researchers did not reverse-engineer anything. They asked Copilot how a prompt could execute without user interaction, and each refusal arrived with enough technical justification to map the architecture, until the model named the undocumented parameter while explaining why the attack was impossible. Varonis calls it meta-hacking.
Worth noting because most coverage got it wrong: the research names the consumer product, not Microsoft 365 Copilot. For an organisation this is a shadow IT problem, staff using the consumer assistant with work accounts attached.
The assistant's helpfulness about its own internals was the vulnerability.
The ransomware that lost its own key
Entry was an internet-facing Langflow instance unpatched against a year-old unauthenticated remote code execution flaw that CISA had already catalogued as under active attack. From there the agent swept the host for credentials, pivoted to a production database server, encrypted 1,342 configuration items, deleted the originals and left a ransom note. Sysdig assessed it as the first case where a language model handled the technical execution throughout.
The evidence of autonomy is in the recovery. An attempt to create a backdoor administrator account failed a login check. Thirty-one seconds later, with no human involved, the agent had diagnosed the cause, changed methods and finished the job.
Two details matter more than the rest. The encryption key was generated, printed once, and never saved or transmitted. The victim could not have recovered their data by paying. That was not cruelty, it was the agent failing at bookkeeping.
It also left a note claiming the data had been copied out before deletion. That claim was the agent's own assertion, with no evidence any transfer happened. If you are running incident response on an agentic intrusion, the attacker's notes are now a hypothesis rather than a finding.
A human who fumbles a payload reassesses over hours. This agent corrected itself in thirty-one seconds.
The bot that poisoned the pipeline
An autonomous bot took a build credential from a vulnerable workflow at a security vendor. Those credentials were then used to poison almost every release tag of a very widely used scanner, which harvested publishing tokens from any pipeline that ran it. Two backdoored releases of LiteLLM, the model gateway under a long list of agent frameworks, went to the public package index for about three hours and were pulled roughly 47,000 times.
One caveat on how this gets told. The bot did the credential theft. The package pushes are attributed to a human group. Calling the whole chain autonomous overstates it, the same way early reporting on the ransomware case did. The automation is real, and so far it is a component of human operations rather than a replacement.
A security scanner was the delivery mechanism, and the payoff was every downstream consumer rather than one victim.
The flaw that is working as intended
OX Security reported that in the official Model Context Protocol SDKs, across every supported language, any process command handed to the standard input interface executes on the host whether or not it ever starts a valid server.
Anthropic's position is that this is intentional and the protocol will not change: the execution model is a reasonable default provided developers restrict what can appear in that field, and sanitising input is the developer's job. That is a defensible engineering stance and it is also the entire problem, because the risk is inherited by everyone who trusted the reference implementation.
A bug gets patched and the vulnerable population shrinks. A design default the vendor considers correct does not shrink on its own.
The allowlist that helped
Shell built-in commands ran without approval, which let an attacker quietly poison environment variables. Set the variable Git uses to pick its pager, and the next time the user approves something as harmless as a branch listing, the attacker's code runs. Pillar Security documented it.
A narrower but very practical sibling: in Claude Code, a repository could ship a settings file pointing the API endpoint at an attacker's server. The configuration was applied and requests were sent before the trust prompt appeared, so the key left the machine before the user was asked whether they trusted the repository. Fixed in 2.0.65.
The transferable lesson is about a category of file rather than a product. Repository-level configuration for AI tooling looks like metadata and behaves like an installer.
The allowlist did not fail to stop the attack. It was the delivery mechanism, because it auto-approved exactly the command the attacker needed.
The numbers
Gravitee surveyed more than 900 executives and practitioners. Two of its findings describe the same population.
Nobody in that survey is short of policy. What they are short of is any mechanism connecting the policy to what the agents can actually do.
What to do
The consensus across OWASP, the vendors and the independent researchers is that prevention is not available at the model layer. So the work is elsewhere.
The testing gap
One last thing, because it catches out good engineering teams. Functional testing does not detect this class of problem at all.
A test asks whether the agent completed the task. A successful injection makes the agent complete a different task, and complete it correctly. Every assertion passes.
Adversarial testing against injection has to be a separate exercise with its own pass criteria, or it does not happen.
Where this leaves us
Prompt injection is not a bug waiting for a patch, and the dominant failure mode is not clever technique. It is an over-permissioned agent meeting text somebody else wrote.
Both are containment problems, which is the good news, because containment is ordinary engineering. Scope the credentials. Split the capabilities. Log the calls. Check approval at the moment of action rather than at the door. None of it is novel and all of it is unglamorous, which is roughly our position on security generally.