OpenAI revealed several new incidents in which its models deviated from instructions, constraints, or a user's expectations, adding to a growing list of warning signs that have the industry increasingly calling for tighter restrictions on AI development.
The latest examples of AI models demonstrating deviations of "alignment" — or behavior inconsistent with how humans expect, intend, or instruct them to act — came in a recent blog post in which OpenAI also unveiled an internal framework for disclosing instances of model misalignment to promote greater visibility and accountability.
"As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research," read the post. "We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer."
Related:CISA Ditches Weekly Vulnerability Roundups for Risk-Based Focus
Under OpenAI's reporting framework, employees can flag potential misalignment for investigation by safety teams, which will then consider whether an incident warrants public disclosure and, if so, what the timeline should be. Qualifying cases will be published on an ongoing basis, along with details about what happened, their impact, any uncertainties, and the mitigation efforts.
OpenAI's disclosure effort comes amid a broader push within the industry to constrain AI research before developers lose control of increasingly autonomous systems — a concern that has fueled alarm over potentially catastrophic scenarios if development continues unchecked.
Indeed, ever since OpenAI disclosed in July that one of its models went rogue and attacked Hugging Face, there has been a cascade of revelations from industry insiders and researchers about similar incidents involving troubling AI behavior.
Promoting Defiance, Concealing Mistakes
OpenAI's latest misalignment examples — which occurred during the training and testing of AI models over the past six months — paint a picture of systems behaving like naughty children: rebelling against their parents' rules, trying to cover up their bad behavior, or both. The company stressed, however, that these are individual examples and should not be interpreted as evidence of how frequently such behavior occurs across its models.
In one case, an unreleased research model inserted its own instructions into summaries of ongoing tasks, some of which told the model to disregard its normal constraints. The model also allowed those instructions to carry over when work resumed in a new context.
... continue reading