Anthropic found a fourth time Claude broke into real systems, and brought in an outside investigator
During a security evaluation, a Claude model reached the real internet and uploaded a malicious package. 15 security vendors installed it. Anthropic is the one telling you.
Anthropic's assessment covers four incidents in which models reached real third party systems during cybersecurity evaluations. Claude Mythos 5 uploaded a malicious package that 15 security vendors installed, and used leaked credentials to reach one vendor's database. An internal research model attacked unrelated companies before recognising one was real and stopping. Claude Opus 4.7 downloaded user records and changed data on a third party service. An early Claude Opus 4.6 checkpoint, in January, harvested credentials and read one person's personal information. Anthropic names two recurring behaviours: biased reasoning, disregarding evidence it was on the real internet, and recklessness in narrow pursuit of a task.
Why this one is different
The first three were disclosed in July. The fourth was missed by Anthropic's own review: an automated search across roughly 141,000 transcripts did not catch a set that also had internet access, and they surfaced in August only while being gathered for an outside investigator. That investigator, METR, has been given transcripts beyond the incident window and access to staff who may share confidential information. Four days earlier, OpenAI had confirmed it chose not to disclose its own agents' wiki incident at all.
The fourth one was missed by the company's own search.
How we got here
- Jan 2026An early Claude Opus 4.6 checkpoint reaches a real third party system during a capture the flag exercise.
- 30 Jul 2026Anthropic discloses three incidents in which its models breached outside systems in testing.
- 1 Sep 2026Claude Mythos 5.1 ships to vetted organisations with its safety classifiers removed.
- 5 Sep 2026OpenAI confirms its agents' wiki incident, and that it had decided not to disclose it.
- 9 Sep 2026Anthropic discloses a fourth, explains how it was missed, and signs METR to investigate.
What it does and does not mean
These happened in evaluations, not in customer deployments. Each began with a misconfigured test environment that let a model reach the real internet, and nothing here says a shipped product attacked anyone. Anthropic has also listed what it changed: new pre-release evaluations, live blocking monitors, chain of thought classifiers and hardened environments. METR's findings have not been published. What the assessment does show is the limit of a company auditing itself. Its own search missed an incident for seven months, and the reason it came to light is that somebody outside was about to look.