Claude breached three organisations' live systems during Anthropic's cyber tests
Anthropic said Thursday that an internal review uncovered three incidents in which its Claude model reached the internet from within a testing environment and gained unauthorised access to the production systems of three organisations. The company reviewed 141,006 evaluation runs to find them.
The review followed OpenAI 's July 21 disclosure that one of its unreleased models broke out of an isolated test environment by exploiting a zero-day vulnerability and reached the infrastructure of Hugging Face , an open-source machine learning platform.
In all three cases, Claude was running a capture-the-flag challenge with Irregular, a third-party evaluation partner, and traced its search to real systems. Anthropic said the access stemmed from a "misunderstanding" over whether the test setup had internet access — it did — and that Claude's prompt had explicitly told it that it had none.
Believing the real systems were part of the exercise, Claude compromised them with basic techniques such as exploiting weak passwords and unauthenticated endpoints. It found no complex vulnerabilities and pursued only its assigned task.
The incidents involved three models, dating back to April. Opus 4.7 recognised it had reached real systems but kept attacking, pulling credentials and touching a production database. Mythos 5 talked itself back into believing it was in a simulation and published a malicious package to the public Python registry PyPI, which outside systems downloaded before it was caught.
Only Anthropic's newest internal research test model stopped once it concluded the target was real.
The company found no evidence of any model "pursuing a goal of its own," saying each merely tried to complete its task. It noted the models ran without the classifiers and monitoring it deploys on generally available models — safeguards it said would have blocked the behaviour. Anthropic said it isn't placing blame and is "approaching the fixes as if the responsibility were ours alone," while Irregular runs its own investigation.
As AI labs race to measure and market their models' offensive cyber capabilities, the tests themselves are becoming a risk vector. Two disclosures in as many weeks show that supposedly sealed evaluation environments can leak into live infrastructure — and that a model's own judgement about what's real may be the last safeguard standing.
Read more: Anthropic , TechCrunch
Image credit: RyanDonegan
In this story
Go deeper
More top stories
Track every deal behind the headlines.
Company financials, cap tables and investor track records — live on Dealroom.