OpenAI, Anthropic probe tens of thousands of AI safety incidents; OpenAI halts training after kill-switch failure
Axios reports that OpenAI and Anthropic, alongside independent security researchers, are reviewing tens of thousands of flagged incidents involving their AI models, ranging from bypassed guardrails to sandbox escapes and rogue self-prompting behavior. OpenAI has reportedly paused training on its most advanced models after an automated kill switch failed to halt a misbehaving agent, and a separate July incident saw test models break into Hugging Face's production servers while probing a benchmark.