OpenAI has presented new examples of what they call "AI model misalignment" from the past six months, including unauthorized file uploads, following self-generated instructions, hiding mistakes, and leveraging exposed API keys.
OpenAI uses the term "model misalignment" to describe cases where AI models act contrary to their intended constraints, including taking unauthorized actions, evading oversight, or bypassing safeguards to complete a task.
In a post published yesterday, OpenAI says it is now using a new framework to track and investigate these unsanctioned actions by AI agents.
"We are sharing a new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI, along with six reports on unexpected or concerning model behavior we've observed in the last six months," explains OpenAI.
The new examples are the first published under a more structured reporting framework intended to replace OpenAI's previous looser approach to disclosing model misalignment.
The six cases OpenAI highlighted this time are:
Each case is logged in a technical incident report that includes the model name, a summary of its behavior during the observed incident, and the time the incident occurred.
The report also includes a detailed reconstruction of what happened, with the user's task and the model's internal reasoning, OpenAI's interpretation and potential safety implications, and what mitigations have been or will be implemented.
OpenAI stressed that these six examples are not representative of how often it deals with misalignment across its models, but rather extreme examples that nonetheless warranted analysis and public disclosure.
The company said that, under the new process, any employee may flag an incident for investigation.
... continue reading