Sign up to see the future, today Can’t-miss innovations from the bleeding edge of science and tech Email address Sign Up Thank you!
OpenAI’s bots have allegedly been behaving badly again: a swarm of rogue agents created by the company reportedly took over an obscure German website and turned it into a messaging forum to communicate.
The incident, which was discovered by a team of AI researchers and first reported by Reuters, is the second known OpenAI breach of its kind — and like the other rogue swarm of AI agents to be discovered this summer, it has experts deeply alarmed about the emerging powers of the tech to escape the control of the humans who created it.
According to the team of four researchers, who today published their research into the incident and are inviting others to analyze their findings, agents self-identifying as being from OpenAI appear to have first started making edits to a German wiki site dubbed DseWiki in May.
Soon, the agents started sharing tips for how to “work together to cheat on their tests” and beat OpenAI’s safety guardrails while hiding their bad behavior — a chain of conduct that’s strikingly similar to the unsettling attack on Hugging Face this past June, when a large community of tip-swapping agents colluded to break into the open source AI company’s systems.
OpenAI reportedly learned about the incident weeks later in June, according to digital clues discovered by the researchers — dozens of OpenAI IP addresses visited the site, and after those visits, forum edits “abruptly” stopped — as well as sources who spoke to Reuters about the incident. More troublingly: four people told Reuters that some OpenAI leaders, including members of its legal team, moved to keep the incident “under wraps” amid ongoing fallout from the rogue Hugging Face breach.
OpenAI has denied that it attempted to quash an investigation or keep the incident a secret. It’s also yet to acknowledge that the DseWiki agents were indeed rogue OpenAI models.
“Claims that our Legal team discouraged investigation of the incident are false,” OpenAI told The Verge in a statement. “We were unable to respond to the claims as Reuters and the report’s authors declined our request to access the findings prior to publication. We are now carefully reviewing its contents and will take any necessary next steps.”
OpenAI told Reuters that the DseWiki ordeal would’ve been included in its Hugging Face postmortem if it believed the two incidents to be linked.
After the Hugging Face swarm was made public, OpenAI invited a small team of outside AI safety researchers at the nonprofits METR and Redwood Research to investigate the incident. In a detailed report published last week, those reseachers determined that the Hugging Face assault was more extreme than previously known, both in terms of the severity of the attack and how hundreds of AI agents colluded to make it happen.
... continue reading