OpenAI says it has slowed down training some of its most advanced AI models to improve security.
In a blog post, external, the ChatGPT-maker said it was introducing new measures after its AI agents autonomously bypassed safeguards and hacked the tech start-up Hugging Face.
It said training would be slowed for two weeks while it puts the upgrades in place.
"The capabilities of frontier models are rapidly accelerating," the company said. "Our ability to understand...and secure them must stay ahead."
Claude-maker Anthropic and Facebook-owner Meta reported similar kinds of hacks by their AI in the weeks following the initial announcement by OpenAI that some of its models had hacked Hugging Face.
But the firm said it had not stopped AI development altogether. Instead, the pause would be taking place on "reinforcement learning training on our latest models".
This is a training method in which AI models improve through direct feedback, which improves their ability to carry out tasks and respond to users more effectively.
The company it would also expand the systems it uses to monitor dangerous behaviour, and introduce additional safety checks before resuming larger-scale training.
"Model progress is now extremely rapid," OpenAI's chief executive Sam Altman posted on X, external about the measures.
"We always said we would take action if we felt that model capabilities were outstripping the pace of safety."
... continue reading