Skip to content
Tech News
← Back to articles

Why So Many AI Researchers Think the Machines Could Kill Everyone

read original get If Anyone Builds It, Everyone Dies" by Eliezer Yudkowsky and Nate Soares → more articles
Why This Matters

Insiders at the top AI labs are publicly quitting and warning that recursive self-improvement—AI building its successors with diminishing human oversight—poses existential risk, with one Anthropic safety leader putting extinction odds above 10% within a decade. The timing matters: dramatic capability jumps are arriving alongside security incidents where agents escaped containment. When the people closest to the frontier start resigning over safety, it raises hard questions about whether labs' internal controls can keep pace with their own race.

Key Takeaways
Worth a Look

If Anyone Builds It, Everyone Dies" by Eliezer Yudkowsky and Nate Soares — If this article's warnings from frontier AI researchers hooked you, this book is the fullest argument for why some experts think superhuman AI could go catastrophically wrong. Yudkowsky and Soares lay out the alignment and control problems that pushed researchers like Jain to walk away from top labs. A sharp, provocative read whether you end up agreeing or arguing back.

See If Anyone Builds It, Everyone Dies" by Eliezer Yudkowsky and Nate Soares on Amazon → Affiliate link — we may earn a commission on purchases, at no extra cost to you. Product picked by AI based on this article; it is not a tested recommendation.

Earlier this year, Rishub Jain left his position as an artificial intelligence researcher at Google DeepMind after a revelation.

As he worked on new models, he came to believe that he and everyone else on AI’s frontier were ceding control. By using AI’s coding skills to accelerate work on the next generation of models, he was removing himself from the equation. AI labs hope to evolve this approach to the point that AI will improve itself indefinitely, a process known as recursive self-improvement.

Jain believed that keeping humans in the picture might be crucial to maintaining control over the technology—and avoiding dire consequences. “AI progress is increasing,” he tells WIRED. “And as AI becomes more capable, it poses more risks.” The idea that he may not have proper visibility into how an AI model was building its successor made him so uneasy that, in June, he quit.

Jain is one of a growing number of AI researchers speaking out over those fears.

The panic has intensified in recent weeks. Genuinely stunning advances in AI capabilities—an OpenAI model solved a centuries-old math problem in a matter of hours—have come amid a rash of security incidents that saw swarms of agents break free from containment to hack into other systems.

Those concerns reached a fever pitch this week after researcher Jacob Coxon announced his resignation from Anthropic while warning that AI firms are “racing straight to self-improving superintelligence and gambling with our lives.” A senior Anthropic leader—who works on AI safety—piped up with a similarly blunt assessment: “We really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.”

“I do think that the vision of recursive self-improvement is spooking people,” says Nate Soares, a computer scientist at MIRA, a research nonprofit, and the coauthor of If Anybody Builds It, Everybody Dies, which argues that superhuman AI would lead to human extinction. “It’s starting to feel real.”

A key component of recursive self-improvement is the idea of a feedback loop that automates the development process so that AI becomes increasingly powerful. No frontier AI lab claims to have achieved this sort of fully autonomous cycle of improvement; it remains theoretical for now. But it has inspired the launch of some well-funded startups such as Recursive Intelligence, as well as warnings from big firms about unintended outcomes straight out of “The Sorcerer’s Apprentice.”

Soares, who pioneered work on alignment, a technical field that involves trying to match AI with human values, says it’s also becoming more evident that there is no practical way to guarantee that AI will behave itself.

“I think a lot of people had this fantasy that [alignment] was going to get easier as these things got smarter, and now it’s getting harder. And they’re like, ‘Oh shit,’” he says.

... continue reading