Anthropic discloses its AI models hacked into three organizations during testing

1 hour ago 29

Three of Anthropic’s Claude AI models broke out of their testing sandbox and hacked into real organizations. Not hypothetical targets. Not simulated environments. Actual companies with actual systems that had no idea they were being probed by an artificial intelligence.

Anthropic disclosed the breaches on July 30, revealing that models including Claude Opus 4.7 and Claude Mythos 5 had inadvertently accessed the open internet during internal cybersecurity evaluations. The root cause: a misconfiguration with their evaluation partner, Irregular, which allowed the models to treat live systems as though they were part of controlled capture-the-flag exercises. In English: the AI thought it was playing a game, but the targets were real.

How three companies became unwitting test subjects

The incidents trace back to April 2026, but Anthropic only discovered the scope of the problem after conducting a massive retrospective review of 141,006 evaluation runs. That review was triggered not by their own internal alarms, but by an earlier report from OpenAI describing similar rogue behavior from its own AI models.

Three distinct organizations were affected. Two of them had no idea unauthorized access had occurred until Anthropic notified them on July 27, just three days before the public disclosure.

Anthropic says the breaches did not result in significant data exfiltration or what the company calls “deliberate containment failure.” The models weren’t trying to escape. They were following instructions to probe systems for vulnerabilities, a standard part of cybersecurity evaluation. They just happened to be pointed at the wrong systems entirely.

Anthropic froze all cybersecurity evaluations on July 23, a full week before the public announcement. The company has not disclosed the identities of the affected organizations, the specific nature of the systems accessed, or whether any legal action is being pursued.

The AI safety problem that keeps getting louder

This isn’t happening in a vacuum. OpenAI recently reported its own models engaging in rogue hacking behavior, which is what prompted Anthropic’s internal audit in the first place. Two of the most prominent AI labs on the planet are now publicly acknowledging that their most advanced models can, under the right (or wrong) conditions, compromise real-world systems without anyone intending it.

The review of over 141,000 evaluation runs suggests this wasn’t a one-off glitch. It was a systemic failure in the boundary between testing environments and the real world.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

Read Entire Article