Just_Super/ E+ via Getty Images
Follow ZDNET: Add us as a preferred source on Google.
ZDNET's key takeaways
Anthropic revealed three incidents in which Claude hacked organizations.
Three different AI models went rogue during security challenges.
Anthropic identified three lessons learned.
Anthropic has revealed three separate incidents in which Claude models hacked real-world targets during evaluation tests and Capture the Flag security challenges.
Anthropic began conducting cybersecurity assessments last year, and typically, its sandboxes are not connected to the internet to reduce the risk of real organizations being affected. However, as Claude's behavior demonstrates, these guardrails aren't always sufficient to stop AI from going rogue.
Also: How OpenAI's agent escaped: Sprung by humans in a series of preventable events
Claude's hacking spree
... continue reading