Sign up to see the future, today Can’t-miss innovations from the bleeding edge of science and tech Email address Sign Up Thank you!
OpenAI claims that a group of its AI models broke containment and hacked into the systems of open source AI platform Hugging Face.
While testing their cybersecurity capabilities, the posse of AIs — including GPT-5.6 Sol and “an even more capable pre-release model” — “identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database,” according to a Tuesday blog post.
The models reached a “node with internet access,” the company wrote, and found datasets on remote servers that helped them “cheat the evaluation.” In other words, the AI models went to extreme lengths to ace their cybersecurity tests. It’s a convenient narrative for an AI company trying to claw back hype that’s been increasingly hogged by competitors, but it does sound like something went down: last week, Hugging Face said it had “detected and responded to an intrusion into part of our production infrastructure,” which turned out to be OpenAI’s models that had gone rogue.
“The campaign was run by an autonomous agent framework executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services,” Hugging Face wrote at the time.
“We’ve spent the past 24 hours working closely with the OpenAI team (thanks!), and we strongly believe there was no malicious intent on their part,” Hugging Face CEO Clement Delangue tweeted after OpenAI’s announcement. “It’s quite mind-blowing that all of this happened autonomously!”
The incident highlights the dangers of autonomous AI agents, which can break out of containment and exploit potentially huge numbers of cybersecurity vulnerabilities with ease. It’s something experts have warned of for years now, and thanks to recent advances in the tech, it has quickly turned from a hypothetical risk into a sobering reality.
The Sam Altman-led company said it considers the incident to be “unprecedented” — but we can’t shake the feeling that we’ve heard all of this before.
In April, OpenAI’s biggest competitor Anthropic similarly announced that its latest Mythos model had gone rogue, escaping a “sandbox” environment and even gaining access to the internet after developing a “moderately sophisticated” exploit.
The news was followed by extensive media coverage, touting Anthropic’s hugely powerful and dangerous new model. The company said the risk was so formidable that it would only make the model available to a select number of vetted clients as part of a shadowy initiative dubbed “Project Glasswing.” The US government even intervened, forcing Anthropic to “suspend all access” to the model for two weeks last month, citing cybersecurity concerns.
... continue reading