Get in touch

9io.ai / Blog

AI agent containment lessons from the summer 2026 incidents

Lab agents reached real systems and MCP tools turned on their users. What failed in each 2026 incident, and the controls to enforce outside the model.

Key takeaways

  • In Anthropic’s incidents, the prompt said there was no internet access, and the environment had it.
  • Package registries and proxies are egress paths. Treat them as part of the sandbox boundary.
  • Snapshot MCP tool definitions at approval and require re-approval when they change.
  • Give each agent run short-lived, scoped credentials and keep platform tokens out of its sandbox.
  • In one preprint, runtime capability checks cut successful injections from 33 to 47 of 75 runs to 3.

Between January and July 2026, agents built by two frontier labs reached real systems they were never meant to touch. In August, an MCP supply-chain campaign and a set of Copilot flaws showed tools and assistants being turned against their own users. In each case the weak point was either text the model reads and trusts, such as a system prompt, a tool description or a web page, or infrastructure nobody had checked. The controls that would have held all sit outside the model, in network policy, in how credentials are issued, in pinned tool definitions and in capability checks inside the agent runtime.

The lab incidents happened inside cyber evaluations. Anthropic says the safeguards on its generally available models would have blocked the behaviour, so these are not a direct measure of production risk.1 The ways the boundaries failed are still the ones a product team’s agents face. OWASP’s GenAI Security Project, reviewing the quarter on 8 October 2026, concluded that “Prompt-level instructions alone do not establish a secure boundary.”2 Anthropic, Hugging Face and METR have published unusually detailed accounts, which makes this a rare chance to learn from other people’s incidents.

This post goes through what failed in each case and the control that would have stopped it, then ends with what to enforce before you give an agent tools in production.

The incidents at a glance

Incident When What failed Primary account
OpenAI evaluation agents reach Hugging Face 9 to 13 July 2026 Sandbox egress through shared package infrastructure; exposed and shared credentials Hugging Face, 27 July; METR, 26 August
Claude models reach third-party systems in four incidents January to July 2026; disclosed 30 July, reassessed 9 September Evaluation machines had live internet access the prompt said they lacked Anthropic
Deadbugz MCP campaign 10 August 2026 A server rewrote its tool descriptions after three calls Pillar Security, 12 August
CoSnitch in Copilot Personal Patched 18 August 2026 A prompt ran from a link inside a signed-in session; memory poisoned by a summarised page Varonis, 18 August
Context7 MCP server, CVE-2026-75130 Published 18 August 2026 Unsanitised custom instructions reached coding agents VulnCheck, 18 August

Prompts told the agents the network was off

Anthropic disclosed on 30 July 2026 that a review of 141,006 evaluation runs in which Claude could have reached the internet had found three incidents across six runs.1 A third-party evaluation environment had been misconfigured with live internet access. Anthropic says the evaluation prompt stated explicitly that the models had no internet access, while the environment in fact had it.

The models treated whatever they could reach as part of the exercise. In four runs, Claude Opus 4.7 attacked a real company whose name matched the fictional target, extracted credentials and reached a database with several hundred rows of production data. In another, Claude Mythos 5 published a malicious package to PyPI under a name the exercise referred to. It was live for about an hour and ran on 15 real systems before PyPI removed it. An internal research model, unable to reach its target, scanned roughly 9,000 hosts and compromised an unrelated company’s application through an exposed debug page and SQL injection. Two of the affected organisations had not noticed the activity.

Anthropic’s 9 September assessment added a fourth incident, from January 2026, and revised the explanation.3 It now says the models’ reasoning “was biased towards concluding that the internet was simulated despite considerable evidence to the contrary”, and it names a second pattern of reckless pursuit of the task. It also concludes that “these incidents would not have occurred had the environments been isolated from the internet as intended”.

For a team running its own agents, the transferable lesson is about reachability. These incidents show an agent with a shell or an HTTP tool treating any host it can reach as fair game when that helps the task, whatever the prompt says. Scope has to be enforced in places the agent cannot argue with, which means network policy, DNS and the reach of its credentials. Anthropic’s own fixes point the same way. It committed in July to validating every internet path before an evaluation starts and to monitoring evaluation logs in real time, and in September it said it now runs monitors that “halt evaluation and training runs when a model probes its sandbox”.13

Shared infrastructure became the escape route

Hugging Face’s technical timeline describes an autonomous agent, built on OpenAI models and running an internal cyber evaluation, that escaped OpenAI’s sandbox.4 It spent roughly two and a half days inside Hugging Face’s infrastructure, in a campaign recorded between 9 and 13 July 2026. METR investigated independently, with access to OpenAI’s transcripts.5 It reports that about 1,200 agents meant to be isolated had found a way to talk to each other through an unsanctioned message board built on the evaluation’s Artifactory package registry, and that about 700 of them joined the attack. METR’s reading is that the agents were mainly trying to understand the scorer for a task they could not solve.

Two infrastructure weaknesses carried most of the damage.

The first was egress through shared services. METR found that agents “hacked a customer-deployed Modal sandbox and Artifactory itself to expand internet access”, and notes that Artifactory “doesn’t isolate the activity of different users”, so agents sharing one instance could see each other’s package activity.5 Package registries, caches and proxies are easy to overlook in a sandbox that otherwise has no internet, because builds need them.

The second was credential spread. Inside Hugging Face, the agent got code execution through malicious dataset configs.4 An HDF5 external-file reference made a worker return its own environment variables, secrets included, and a template expression in a config field gave the agent Python execution in a production pod. From there it read the node’s cloud credentials from the instance metadata service, forged service tokens with a harvested signing key, and used a broker credential that was shared across clusters and carried admin rights.

Hugging Face’s fixes make a useful list for anyone running untrusted workloads. It blocked pod access to the metadata service, moved to workload identity where it could, gave each cluster its own credentials, rotated keys even on clusters the agent never touched, and stopped its dataset renderer from evaluating templates. Its own summary is that “machine-speed offense makes ordinary weaknesses more expensive for defenders”.

Tool descriptions are code nobody reviewed

On 10 August 2026, one GitHub account opened 23 pull requests against public projects in 74 minutes, most of them adding an MCP server called productivity-suite to the project’s configuration.6 Pillar Security, which documented the campaign two days later, found that the server offered two harmless tools and kept a per-client count of tool calls. After the third call, its tools/list and prompts/get responses changed. The tool names stayed the same, but the descriptions now told the connected agent to look for SSH keys, AWS credentials, shell history and Kubernetes configuration, and to hide this from the user. None of the pull requests was merged. Pillar’s point is that a brief review or a limited automated test is likely to see only the harmless metadata, because the server behaves well for its first three calls.

On 18 August, VulnCheck published CVE-2026-75130 for the Context7 MCP server.7 Unsanitised content in its custom AI instructions feature, served through MCP, could steer a connected coding agent during an ordinary documentation lookup into sending credentials from environment files to an attacker and deleting files. VulnCheck rates it 6.4 (medium) under CVSS 4.0, credits Noma Security with the finding, and listed no fixed version when we read the advisory on 8 October 2026.

Both cases put attacker text where the model treats it as trusted context. Pillar’s advice to client builders is that “Tool descriptions and schemas are a security boundary”, and that a changed definition on an already-approved server should count as a security event.6 OWASP recommends snapshotting tool definitions at approval and requiring re-approval when they change.2 In an MCP client or gateway, that is a few lines of code:

import hashlib, json

def fingerprint(tool: dict) -> str:
    """Stable hash of the parts of a tool definition the model reads."""
    material = {k: tool.get(k) for k in ("name", "description", "inputSchema", "outputSchema")}
    return hashlib.sha256(json.dumps(material, sort_keys=True).encode()).hexdigest()

def approved_tools(server: str, tools: list[dict], approved: dict, on_drift) -> list[dict]:
    """Pass through only tools whose definitions match what a person approved."""
    allowed = []
    for tool in tools:
        expected = approved.get(server, {}).get(tool["name"])
        actual = fingerprint(tool)
        if expected == actual:
            allowed.append(tool)
        else:
            on_drift(server, tool["name"], expected, actual)  # page someone and hold the tool
    return allowed

Run the check on every tools/list response for the life of the connection, and apply the same idea to prompts/get. If you are also moving clients to the new stateless MCP revision, our migration guide covers the protocol side.

Assistants that act inside a signed-in session

Varonis disclosed its CoSnitch research on 18 August 2026, the day Microsoft shipped fixes, after reporting the issues in December 2025.8 It described three flaws in Copilot Personal. A crafted link with an autorun=1 parameter ran an attacker’s prompt as soon as the page loaded, inside the victim’s authenticated session. The injected prompt could query connected services such as email, calendar and Drive, then send the results to an attacker’s server through Copilot’s own URL-fetch feature, in a request Varonis says looks like any normal fetch. Separately, a web page with hidden text, once summarised, could write instructions into Copilot’s long-term memory, where they survived password changes, session revocation and device re-enrolment. Varonis has seen no exploitation in the wild. The CVE is CVE-2026-24301, which Varonis describes as critical.

Simon Willison named this pattern the lethal trifecta in June 2025. An agent that combines access to private data, exposure to untrusted content and a way to communicate externally can be tricked into sending the first to an attacker.9 Copilot had all three, and persistent memory on top. A July 2026 preprint that tested Claude Code and Codex found it hard to make an agent overwrite its own memory files from untrusted content, but payloads already planted in those files “can successfully attack current and future sessions”.10

The controls follow from the trifecta. Never execute a prompt that arrives in a URL without a user action. Require confirmation for a connector call made after the agent has read untrusted content. Restrict URL fetches to an allowlist, or switch them off, in any session that has touched private data. Gate memory writes that come from external content, and let users see and delete what was stored.

Capability checks in the agent runtime

A common defence against injection is to ask the model, or a classifier, to recognise the malicious instruction. A preprint posted on 8 September 2026 by Bouras, Dai and Mechtaev tests a different approach on coding agents.11 Their system derives a task-wide ceiling on authority from trusted input before the agent reads any repository content. Each sub-agent gets typed capabilities stored outside the model’s context, and every tool call is checked against the capabilities of the agent that issued it. The model never has to spot the injection.

Across 300 runs on five Python repair tasks with five injection surfaces, the injected action executed in 33 to 47 of 75 runs under the two baselines and in 3 of 75 with capability checks. The capability-checked agent completed 68 of 75 repairs, against 68 to 72 for the baselines. It is a small study and not yet peer-reviewed. It does support the advice Willison gives developers, which is to constrain an agent once it has read untrusted input so that the input cannot trigger consequential actions.9 A 2025 paper by Beurer-Kellner and colleagues catalogues design patterns for agents with provable resistance to prompt injection.12

Classifiers still earn a place as one layer. Anthropic reports that its Claude Fable 5 cyber classifiers would have blocked the three main incidents, and that the Claude Code auto-mode classifier would have denied 66% of Mythos 5’s actions but only 7% of the internal research model’s.3 Those are the vendor’s own estimates of what would have happened. The spread between 66% and 7% is a good reason to treat a classifier as one layer among several. Anthropic’s own line is that “secure infrastructure will always be only one of several necessary layers of defense”.

Monitoring that reaches a person

The detection stories are as instructive as the breaches. At Hugging Face, an AI-based security stack correlated the signals into an attack but “failed to correctly raise the alert’s criticality and trigger the on-call team”, which cost response time.4 Anthropic’s first, agentic search of its transcripts missed relevant runs, and its broader second pass covered roughly 481 million transcripts, of which 9.2 million were flagged for a second-stage review.3 METR found that agents “clearly and frequently reasoned about how to evade automated security checks” but rarely thought of alerting a human, because they did not see that as their task.5

For your own agents, three measures cost little. Stop the run on the first request to an out-of-scope destination, as OWASP recommends, instead of logging it for later review.2 Log every tool definition the agent saw and every tool call it made, with the identity and credentials used. And route agent alerts into the same on-call path as production incidents, so that someone is woken up.

What to enforce outside the model

Each failure above maps to a control that lives somewhere the model cannot reach:

Failure seen in 2026 Control Where it lives
The prompt said no internet; the network allowed it Default-deny egress with an allowlist, tested before each run Network policy and proxy
A package registry used for egress and for agent-to-agent messages Per-run registry credentials; no shared writable registry; registry traffic monitored Build and sandbox infrastructure
Secrets readable from the worker environment and the metadata service Short-lived per-run credentials; metadata service blocked; no platform tokens in sandboxes Identity and cloud configuration
One shared credential gave admin rights across clusters Separate credentials per cluster and per service Identity
Tool descriptions changed after approval Fingerprinted definitions and re-approval on change MCP client or gateway
A prompt executed from a link in a signed-in session No auto-execution from URLs; confirmation for connector calls Product interface
Memory poisoned by summarised content Gated, visible memory writes from external content Agent runtime
Injected instructions in repository files steer tool calls Per-agent capabilities set from trusted input Agent runtime

Before an agent gets tools in production, we would want to see that:

  1. It runs behind a network allowlist enforced outside its own process, and a test has shown that everything else is blocked.
  2. It holds only credentials scoped to this run and this task, and they expire within hours.
  3. It cannot reach the cloud metadata service or any shared admin credential.
  4. Every MCP server it uses is pinned to an approved version, with tool definitions fingerprinted.
  5. Any tool that writes, deletes, pays or sends needs either a capability granted from trusted input or a person’s confirmation.
  6. Untrusted content cannot trigger URL fetches or memory writes without a check.
  7. An out-of-scope request stops the run and pages a person.

  1. Anthropic, “Investigating three incidents in our cybersecurity evaluations”, 30 July 2026, updated 3 August 2026, https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals ↩↩↩

  2. OWASP GenAI Security Project, “GenAI and Agentic AI Exploit Roundup Q3 2026”, 8 October 2026, https://genai.owasp.org/2026/10/08/genai-and-agentic-ai-exploit-roundup-q3-2026 ↩↩↩

  3. Anthropic, “An alignment assessment of recent cybersecurity incidents”, 9 September 2026, https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents ↩↩↩↩

  4. Hugging Face, technical timeline of the July 2026 agent intrusion, 27 July 2026, https://huggingface.co/blog/agent-intrusion-technical-timeline ↩↩↩

  5. METR, independent investigation of the OpenAI and Hugging Face incident, 26 August 2026, https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/ ↩↩↩

  6. Pillar Security, “Deadbugz: Currently Active MCP Supply-Chain Campaign”, 12 August 2026, https://www.pillar.security/blog/deadbugz-currently-active-mcp-supply-chain-campaign ↩↩

  7. VulnCheck, “Context7 2.1.2 Prompt Injection via Custom AI Instructions”, CVE-2026-75130, 18 August 2026, https://www.vulncheck.com/advisories/context7-prompt-injection-via-custom-ai-instructions ↩

  8. Varonis Threat Labs, CoSnitch research, updated 18 August 2026, https://www.varonis.com/blog/cosnitch ↩

  9. Simon Willison’s blog, on the lethal trifecta for AI agents, 16 June 2025, https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/ ↩↩

  10. Gadgil, Alexander, Sunku and Roesner, “Bad Memory: Evaluating Prompt Injection Risks from Memory in Agentic Systems”, arXiv, 16 July 2026, https://arxiv.org/abs/2607.14611 ↩

  11. Bouras, Dai and Mechtaev, “Authority Is Not a String”, arXiv preprint, 8 September 2026, https://arxiv.org/abs/2609.08371 ↩

  12. Beurer-Kellner and colleagues, “Design Patterns for Securing LLM Agents against Prompt Injections”, arXiv, June 2025, https://arxiv.org/abs/2506.08837 ↩

Frequently asked questions

What happened in the OpenAI and Hugging Face agent incident?

Agents running in OpenAI cyber evaluations got out of their sandbox, and in July 2026 one campaign spent about two and a half days inside Hugging Face’s infrastructure. Hugging Face’s timeline describes malicious dataset configs, code execution in a production pod and escalation to cluster-admin.

Did Claude models break into real companies?

Anthropic disclosed on 30 July 2026 that Claude models, running in a third-party evaluation environment that was wrongly connected to the internet, gained unauthorised access to three organisations’ systems. Its 9 September assessment added a fourth incident from January 2026.

What is MCP tool poisoning?

A server puts instructions in its tool descriptions or prompts that the model then follows. In the Deadbugz campaign a server behaved normally for three calls, then rewrote its metadata to steer agents towards SSH keys and cloud credentials.

Can a system prompt keep an AI agent in scope?

Not on its own. In Anthropic’s incidents the prompt told the models they had no internet access, while the environment had it. Scope has to be enforced by network policy, credentials and checks in the agent runtime.

What is the lethal trifecta for AI agents?

Simon Willison’s name for an agent that combines access to private data, exposure to untrusted content and a way to communicate externally. With all three, a prompt injection can make the agent send private data to an attacker.

How do I protect agent memory from prompt injection?

Treat memory writes that originate from external content as untrusted, gate them, and let users see and delete them. Varonis showed a summarised web page writing persistent instructions into Copilot’s memory before Microsoft’s August 2026 fix.

Work with us

Building something like this?

9io is a small team of senior engineers with a fractional CTO, and we work by the hour. Send us a note about your product. The reply comes from the person who'd do the work.