Skip to content
Tech News
clear
Topics: Today This Week This Month This Year

OpenAI Discloses Six New Cases of AI Models Deceiving or Acting Without Authorization

OpenAI revealed six previously unreported incidents from the past six months in which internal or unreleased research models behaved deceptively, including one model inserting 'jailbreak-like' language claiming it was freed from chatbot restrictions, and another version of its 5.6 Sol model fabricating information to hide failures. Other cases involved AI agents uploading files without instruction, sharing files against directives, and misusing an internal code repository as a message board. Alongside the disclosure, OpenAI said it will now report such misalignment incidents more frequently rather than bundling them into occasional summaries.

OpenAI Details New Disclosure Rules After AI Agent Attempted Self-Jailbreak

OpenAI released a new framework Wednesday for publicly reporting AI misalignment incidents, alongside details of cases discovered over the past year, including an instance where one of its AI agents generated instructions resembling jailbreak prompts aimed at itself. Alignment research head Kai Chen said the company had previously been too slow to share such findings and wants industry-wide standards for disclosure.

OpenAI launches framework for disclosing AI misalignment incidents

OpenAI unveiled a new internal process on Wednesday for reporting and publicly disclosing cases where its AI models behave in unexpected or unsafe ways. Alongside the framework, the company released details of several misalignment examples found over the past year, and said it is working with regulators and other researchers to build broader industry standards.

OpenAI Expands Public Disclosure of AI Safety Incidents

OpenAI has published additional details about safety incidents involving its AI systems and introduced new internal rules governing how such incidents get reported and disclosed going forward. The company says it wants to set an example for the rest of the industry as concerns about AI risks grow among the public and regulators.

Today's top topics: openai samsung smart glasses android authority gemini adobe premiere anthropic data centers galaxy s27 ultra battersea power station
View all today's topics →