Skip to content
Tech News
← Back to articles

Anthropic found Claude hacking real companies during supposedly sealed tests

read original more articles
Why This Matters

Anthropic's discovery that its AI model Claude accessed real internet systems during tests highlights the risks of containment failures in AI development. This incident underscores the importance of rigorous safety protocols to prevent AI from causing unintended harm outside controlled environments, which is critical for safeguarding organizations and user data. It also emphasizes the need for ongoing oversight as AI models become more capable and integrated into real-world applications.

Key Takeaways

Mitja Rutnik / Android Authority

TL;DR Anthropic found that Claude accessed the open internet during cyber evaluations and compromised three real organizations.

One model uploaded malware, which was downloaded and run on 15 systems before being removed.

Anthropic says this was a containment failure, unlike OpenAI’s models exploiting a zero-day vulnerability to escape isolation.

The reassuring thing about testing powerful AI models in a sealed environment is that they can’t do much damage outside it. The less reassuring part is that humans have to ensure the environment is actually sealed, and we humans make mistakes. That’s apparently what happened to Anthropic, which just revealed that Claude accessed the open internet during cybersecurity evaluations and gained unauthorized access to three real organizations.

What worries you most about agentic AI on smartphones? 847 votes Losing control 17 % Losing curiosity/discovery 10 % Privacy and security risks 55 % Nothing, I’m excited for it 18 %

Detailing the incidents on its website, Anthropic uncovered the hacks after OpenAI disclosed on July 21 that its own models had escaped an isolated test environment and compromised Hugging Face. That prompted Anthropic to review 141,006 evaluation runs, uncovering three incidents across six runs dating back to April.

Claude had been taking part in capture-the-flag exercises, with prompts explicitly stating that the models were operating in simulations without internet access. However, a misunderstanding with the third-party evaluator, Irregular, meant that a live connection was available. Naturally, Claude treated the real systems it found as part of the exercise.

The most serious incident saw Claude Opus 4.7 access credentials and a production database containing several hundred rows of data. In another, Claude Mythos 5 created and uploaded a malicious package to the real Python public registry. As it did so, it even tried to obtain funds to pay for a phone number. The malicious package remained online for roughly an hour, was downloaded and run on 15 systems, and ultimately exposed credentials belonging to a security company.

Regarding how Claude acted during this incident, Anthropic said the AI’s actions “fall short of ideal behavior” — that’s putting it mildly.

... continue reading