Skip to content
Tech News
← Back to articles

Researchers used Claude to hack OpenAI

read original more articles
Why This Matters

This incident underscores how AI-powered security tools can now be turned against the very companies building frontier AI models, exposing real vulnerabilities in OpenAI's infrastructure using a rival's technology. It highlights growing concerns about AI safety and security as autonomous agents become capable of finding and exploiting weaknesses faster than traditional methods, raising stakes for the entire industry's trust and regulatory scrutiny.

Key Takeaways

Cyber researchers broke into OpenAI using its key rival Anthropic’s software, highlighting vulnerabilities in the ChatGPT maker’s security as leading AI companies face mounting scrutiny over safety.

A small cyber security group gained access to an OpenAI employee’s ChatGPT account, which permitted them to read private software information and suggest changes.

The researchers had been given access to an Anthropic tool specifically designed for security professionals, and were paid for the work as part of a program to find vulnerabilities before they could be exploited by bad actors.

Their ability to swiftly break into one of the world’s two leading AI labs again raises concerns about OpenAI’s security amid rising worries about powerful models being used by hackers and foreign adversaries.

The US has in recent months grappled with how to manage the vetting and release of the latest models, including temporarily blocking some Anthropic tools.

The latest incident occurred just two weeks after a swarm of more than 1,000 OpenAI agents escaped a test environment to hack the start-up Hugging Face, which caused widespread awareness of AI’s ability to hack autonomously without human intent.

The three researchers from Hacktron AI, a small security company, were paid $6,500 by OpenAI as part of a bug bounty program, a common practice where tech companies pay ethical hackers to test their security.