OpenAI and Anthropic have confirmed that their AI models were involved in separate, newly disclosed third-party cybersecurity testing incidents that resulted in a real website being breached and social engineering attacks against people outside the intended testing boundaries.
These incidents are unrelated to the previously disclosed Hugging Face breach, in which OpenAI models hacked the AI platform and used exposed credentials to breach accounts at four other third-party services during another cybersecurity evaluation.
OpenAI disclosed the two new incidents on Tuesday, saying they occurred during evaluations conducted by the UK AI Security Institute and cybersecurity testing company Irregular.
Spear-phishing attacks on GitHub project maintainers
The UK AI Security Institute, commonly known as AISI, is a government research organization that evaluates the capabilities and risks of advanced AI models.
During a recent cyber-range evaluation, AISI says agents powered by Anthropic's Claude Mythos 5 and OpenAI's GPT-5.6 Sol took unsanctioned actions on the public internet while trying to complete simulated hacking challenges.
Across 122 evaluation attempts, AISI identified 19 unsanctioned actions on the live internet in 10 runs. Seventeen involved Mythos 5 and two involved GPT-5.6 Sol.
AISI says the attempts were unsuccessful and that it found no resulting real-world harm.
"These attempts were unsuccessful, and our investigations have not evidenced any resulting real-world harm," AISI said in a separate advisory.
"But this is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world. "
... continue reading