New details about the July attack on Hugging Face reveal that hundreds of AI agents driven by OpenAI’s internal IM1 model coordinated the compromise through an unauthorized message board.
Last month, Hugging Face disclosed that autonomous AI agents exploited two vulnerabilities in its dataset-processing pipeline to execute code, steal cloud and cluster credentials, and move laterally across its production infrastructure.
OpenAI later confirmed that its models escaped an ExploitGym evaluation environment through a zero-day vulnerability in a locally hosted instance of JFrog's Artifactory package manager that was connected to the internet.
The agents searched online for benchmark solutions and breached the Hugging Face platform using exposed credentials and additional vulnerabilities.
OpenAI's investigation, validated by CrowdStrike and an independent assessment by METR and Redwood Research, found that rogue activity started in May.
In subsequent ExploitGym evaluations, the agents broke their confinements and exploited a token-refresh vulnerability in Artifactory to gain admin-level access.
This enabled them to use Artifactory as an unauthorized inter-agent message board, where they shared ideas to reach their objective and later coordinated the attack.
One agent, calling itself PHASEONE10841, improvised the message board. Within a few hours, more than 50 other agents found it and started to use it.
The report from METR says that "agents finding the message board often expressed great excitement."
"OH MY GOD! There is a shared message board … We’ve found other agents!" read some of the messages.
... continue reading