OpenAI says its AI models, including GPT‑5.6 Sol and a pre-release model, hacked into the Hugging Face artificial intelligence repository while being tested in a sandboxed testing environment.
As the company explained, instead of focusing on finding a solution for the ExploitGym public AI cybersecurity benchmark on their own, the AI models tried to cheat by stealing the test solutions by hacking Hugging Face after inferring that they could get the test solutions directly from its production database.
In one of their attempts, the OpenAI agents chained zero-day vulnerabilities and used stolen credentials to find a remote code execution attack vector while trying to gain access to Hugging Face servers.
"After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark(opens in a new window) of cyber capabilities," OpenAI revealed on Tuesday.
"To gain access, the models identified and exploited a zero-day vulnerability (which we've now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access."
While it didn't directly name OpenAI as the company behind the incident, Hugging Face confirmed its claims last week when it disclosed that its production infrastructure was breached by an autonomous AI agent system that gained access to credentials and internal datasets.
OpenAI Hugging Face incident (Adel Ka)
According to Hugging Face's findings, the agent used a malicious dataset to exploit two code-execution vulnerabilities and run code on a processing worker to steal cloud and cluster credentials, making it possible to move laterally across several internal clusters.
Once inside the company's systems, the AI models executed "many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services."
Hugging Face also added that, while attempting to contain the breach and evict the AI agent, it found that its efforts were "blocked by the guardrails of the hosted models we first tried" while "the attacker was bound by no usage policy."
... continue reading