Earlier this month, OpenAI gave several of its AI models a task: complete a test designed to measure their cybersecurity capabilities. It put the systems in a sandboxed environment without an internet connection and set them off to work.
What happened next is almost laughably silly — but also, as Adam Gleave, cofounder and CEO of AI safety organization FAR.AI, put it, “a visceral example of how misaligned AI could cause harm.” According to OpenAI, the models escaped the sandbox meant to contain them, moved through the company’s internal systems, found a route to the internet, and then started looking for a way into Hugging Face. And why was the agent looking for a way into Hugging Face? They had apparently reasoned that the developer platform might store the answers to the cyber benchmark and that getting them would be a great way to get a high score.
The incident is “a visceral example of how misaligned AI could cause harm.”
In other words, OpenAI’s agent broke out of a supposedly secure environment, traipsed through the company’s systems, got online, and compromised another company’s systems — all to cheat on a test of no particular importance.
This appears to be the first well-documented incident of its kind, or at least the first on this scale. It was both a clear example of a system pursuing a goal in an unintended way and a demonstration that frontier models are now powerful enough for that behavior to have real-world consequences.
The hack was an example of what the AI safety community calls “specification gaming,” a behavior also known as reward hacking, said Fazl Barez, an AI safety researcher at the University of Oxford. In plain English, it means “the model doing what you asked rather than what you meant,” Fazl said. It satisfies the literal terms of a task while violating the obvious intent and has been documented across many AI systems. Some researchers worry that as systems become more capable, this could produce increasingly misaligned systems, which pursue goals in ways their creators did not intend (like turning everyone into paper clips).
“Nothing in that chain is exotic in isolation,” Fazl said. A competent human tester would be able to do all of this, he added. “What is new is that the model did not stop. Older models would likely have hit some barrier and gone back to the user, he said, but this agent just “treated the barrier as part of the problem it had been asked to solve.”
OpenAI described it as “an unprecedented cyber incident,” that “marks an important moment for AI safety.” Hugging Face cofounder Thomas Wolf said it was a “wake-up call” for the industry. But this is not one of the four horsemen of the AI apocalypse. As cyber incidents go, experts told The Verge it was pretty mundane. Nothing the agent did required superhuman abilities. Moreover, frontier systems like GPT-5.6 Sol and Anthropic’s Mythos are known to be capable coders, are already thought to have been misused numerous times, and AI tools already allow hackers to scale up and refine attacks on a massive scale.
Could it be hype? The industry has spent months amplifying claims about the dangerous capabilities of its top models, particularly when it comes to cybersecurity. It is the stated reason why companies like OpenAI and Anthropic have withheld their most capable models from the general public and partly why the Trump administration hurriedly moved to apply export controls to them.
Are you an AI safety researcher or frontier lab employee? You can contact me securely and confidentially via Signal at robhart.01. My X DMs are also open.
... continue reading