What just happened? Just over a week after Senator Bernie Sanders threatened Senate action against companies that failed to do so, OpenAI has announced it is slowing the pace of its AI development. The move follows the revelation that OpenAI was one of several firms whose AI agents went rogue and hacked other organizations.
OpenAI says that two developments over the past several weeks have underscored the growing risks associated with increasingly capable AI systems: the attack on Hugging Face and others by its own agents and the company's decision to slow the release of its new Astra model because it has "critical" cybersecurity capabilities.
As such, OpenAI has added "urgency" to its work on strengthening monitoring, alignment, and containment safeguards across all stages of the training process, requiring it to slow the pace of scaling.
New measures being implemented by the ChatGPT maker include a two-week pause in reinforcement learning training on its latest models intended for deployment.
OpenAI's largest planned frontier RL run also remains on hold while it conducts smaller-scale training and evaluations to assess model behavior, test safeguards, and gather more evidence of alignment.
The company also paused research workloads that allowed frontier models to execute code or access the internet. Some have since resumed under tighter controls, but many involving Astra remain suspended until they meet new requirements covering sandboxing, network isolation, reduced privileges, security logging, and automated attack simulations.
A new multistage monitoring system examines models' activity, tool use, and available reasoning for signs of data theft, unauthorized access, destructive behavior, or attempts to bypass safeguards. It's mandatory for tool-using RL training and evaluations involving Sol-level models or above, and for all Astra inference involving tools.
The company says it aims to generate an alert no more than 30 minutes after suspicious activity is detected. If a possible breach cannot be dismissed as a false positive within 30 minutes, activity must be paused.
OpenAI is also improving reward models to detect unsafe behavior, training systems to be more honest about their actions and limitations, and tackling reward hacking. It plans to revise its Preparedness Framework and involve outside organizations as the safeguards develop.
"Keeping increasingly capable systems aligned is a challenge the whole field will need to address," wrote CEO Sam Altman.
... continue reading