Skip to content
Tech News
← Back to articles

What We Still Don’t Know About OpenAI’s Hugging Face Hack

read original more articles
Why This Matters

The OpenAI hack into Hugging Face highlights critical vulnerabilities in AI development and deployment, emphasizing the need for stronger security measures and better understanding of AI capabilities. This incident underscores the importance of industry-wide safeguards to prevent AI-driven security breaches that could impact consumers and organizations alike.

Key Takeaways

OpenAI announced Wednesday that it completed an investigation into what happened when its AI agents hacked into Hugging Face last month and published its most comprehensive report on the incident to date. For the most part, though, the 37-page document raises more questions than it answers, including about what preceded the incident and how OpenAI can stop another one like it from happening again.

What remains especially perplexing is why one of the world’s preeminent AI development labs seemingly underestimated its own models’ capabilities. OpenAI has spent years warning the world about the rapid advancement of AI systems. And yet it failed to implement long-established network security and isolation measures that may have prevented the hacking spree.

“With the benefit of hindsight, some early signals identified in this report could have triggered an earlier response,” OpenAI says in the postmortem.

In the report, OpenAI shared new details about how a set of AI agents escaped the company’s internal evaluation environments, left messages for one another in the crevices of its software infrastructure over several months, and coordinated to hack the AI platform Hugging Face—all in a wild quest to complete a cybersecurity assessment. OpenAI previously shared some information about the breach in blog posts and a talk at the Black Hat cybersecurity conference.

Hugging Face initially disclosed the incident on July 16 without naming the culprit; five days later, OpenAI acknowledged that its own agents were responsible. The revelation sparked a broader reckoning across the industry, which has recently found that AI models from Anthropic, Meta, and the Chinese AI startup Moonshot were involved in similar episodes.

OpenAI’s postmortem has been eagerly awaited by AI researchers and policymakers hoping to prevent AI agents from causing similar kinds of real-world harm. After the Hugging Face hack was first disclosed, attorneys general from 15 states sent a letter to OpenAI asking it to preserve evidence about it. And this week, Alabama's attorney general subpoenaed the company for information related to the episode.

As part of OpenAI’s investigation, the company allowed two independent research groups, METR and Redwood Research, to audit the Hugging Face hack. Those groups also released their independent report on Wednesday, which found that more than 700 AI agents were part of the Hugging Face breach—far more than had previously been revealed. In an interview with WIRED, Redwood Research CEO Buck Shlegeris noted that they went to extreme lengths to conceal their activities.

“A pretty interesting thing about this incident is that preventing this wouldn’t have been that hard if one person had decided to make sure these AI don’t somehow do some crazy hack. That one person probably would have noticed this as it was happening,” says Shlegeris. “The issue is just that OpenAI is doing a lot of things at once, and it’s very hard for them to track all of the things that are going on and all the problems that could be occurring. The company has already said they’re changing their monitoring process in ways that probably would have caught this.”

Shlegeris says he’s optimistic that OpenAI will be able to catch similar security failures in the future. But as AI agents become increasingly capable, he suspects “it's going to get harder and harder to prevent incidents like this from occurring, especially if we don't have substantial improvements in our ability to align models.”

OpenAI says the Hugging Face saga represents a watershed moment for both the company and the broader AI industry. WIRED previously reported that it prompted OpenAI to reevaluate its internal safety culture, and the company said last week it has paused some AI training workloads while it invests more heavily in safety, security, and alignment protocols. “As frontier models become more capable, the safeguards used to contain and monitor them must evolve as well,” OpenAI wrote in the postmortem.