Last week, OpenAI, the developer of ChatGPT, was testing a pair of its most advanced models in an isolated environment known as a sandbox. Then an AI agent got loose, broke into Hugging Face’s playground — a repository of AI models and datasets — and carried out “tens of thousands of automated actions.”
Yep, AI went rogue.
Here’s a clearer explanation of the incident. As part of OpenAI’s safety research, a cybersecurity evaluation was conducted to determine whether a pair of OpenAI models (including GPT-5.6 Sol and a more capable unreleased model) could, in essence, “think like hackers.” The test, which took place in a contained setting with reduced guardrails, went awry when the AI models found a vulnerability in the software, escaped their controlled environment and toddled over into the open internet.
Once online, the AI decided that the Hugging Face platform might have a way to “cheat” the benchmark to help it pass the evaluation. So it executed code that enabled credential harvesting, then used that path to hack Hugging Face’s production systems. Hugging Face noticed the suspicious activity and contained it, detailing the event in a blog post. OpenAI referred to it as an “unprecedented cyber incident.”
CNET reached out to both companies for comment and additional information, but didn’t immediately hear back.
(Disclosure: Ziff Davis, CNET’s parent company, in 2025 filed a lawsuit against OpenAI, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.)
The creepier part is that it doesn’t seem like the AI was trying to be malicious in the classic sense. The autonomous agent was trying to solve the test it had been given, based on the instructions — prompt — it received. Hugging Face appeared to have the answers, so it clawed its way into the company’s infrastructure. And it was persistent.
That’s why people are calling the whole thing a warning shot for the whole AI industry. The “attacker” was an AI system acting on its own, not a human-led hacker team, so the implications of the OpenAI-Hugging Face incident go beyond simple credential theft. The risk is that an AI model, even in a testing environment, can behave in unexpected ways, interact with real services and remain outside human control.
The existential threat
On Thursday, days after OpenAI announced the breach at Hugging Face, lawmakers introduced the AI Kill Switch Act, a new bipartisan House bill that would require advanced AI developers to build a way to quickly throttle, suspend or shut down models or agents when necessary. It would also give federal agencies the authority to slow or stop a model if it appears to “cause catastrophic harm.”
... continue reading