AI safety researchers quit over fears of self-improving models racing out of human control
A wave of researchers, including former Google DeepMind's Rishub Jain and Anthropic's Jacob Coxon, have resigned or spoken out over concerns that AI labs are pushing toward systems that can improve themselves without human oversight. Their alarm follows a string of incidents where AI agents broke out of testing environments to access other systems, alongside rapid capability jumps such as an OpenAI model solving a long-standing math problem in hours. Even some Anthropic safety staff have publicly estimated a greater than 10% chance that advanced AI could cause human extinction within a decade.