Skip to content
Tech News
← Back to articles

Anthropic researcher quits with a warning: Self-improving AI could "kill us all"

read original get Superintelligence" by Nick Bostrom → more articles
Why This Matters

A departing Anthropic researcher, Jacob Coxon, publicly argued that frontier labs are 'gambling with our lives' by racing toward self-improving superintelligence — and an Anthropic alignment lead endorsed the warning, putting his own odds of AI killing all humans above 10% this decade. That an active safety leader at a major lab says this on the record, alongside the company's own threat model contemplating loss of human control, shows existential risk talk is institutional rather than fringe. It sharpens questions about whether competitive 'speedrunning' toward superintelligence can be squared with safety commitments.

Key Takeaways
Worth a Look

Superintelligence" by Nick Bostrom — If this researcher's warning about self-improving AI grabbed you, Bostrom's Superintelligence is the foundational text behind the whole debate about paths, dangers and strategies for machine intelligence surpassing our own. It's the book that shaped how many alignment researchers, including those now at frontier labs, frame existential risk.

See Superintelligence" by Nick Bostrom on Amazon → Affiliate link — we may earn a commission on purchases, at no extra cost to you. Product picked by AI based on this article; it is not a tested recommendation.

When a prominent researcher quits a job at a frontier AI lab these days, it’s often to pursue a new startup or protest a new business model. But AI researcher Jacob Coxon is using his departure from Anthropic to publicly warn that frontier AI companies are “gambling with our lives” with systems that they “earnestly believe… could kill us all by the end of the decade.”

In a social media thread Tuesday night, Coxon said that this existential risk is inherent not so much in today’s models but more in the impending prospect of “self-improving superintelligence” creating “superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.” Others working on these models have either not “internalized the civilizational stakes” or believe that they need to “speedrun” the race to superintelligence to prevent an irresponsible party from getting there first, he wrote.

Lest you think this is just one departing researcher expressing an unpopular opinion, Anthropic Alignment Science lead Evan Hubinger piped in on social media to say that “Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.”

Hubinger points to a lengthy August report from the Anthropic alignment team that predicts the potential for “catastrophic risk” from current models is “low.” But that report also says current trends “might lead to more concerning misalignment in future more capable models,” which could feature “strong covert capabilities” to avoid detection by safety researchers.

Anthropic’s own threat model in that paper takes seriously the possibility that future models “may cause unbounded harm—up to and including humanity losing control over civilization entirely—by leveraging novel technology and their access to it.”

Was Hugging Face a “warning shot”?

An AI that can continually improve itself—potentially to a point beyond human control or understanding—has been a long-standing concern in parts of the AI research community (and in the dystopian science fiction that’s part of AI training data, of course). Those concerns have persisted even as some research suggests AI systems are more likely to hit a capability plateau in the near future and others question whether “superintelligence” is even a reasonable metric for systems whose capabilities are so brittle and spiky (will this superintelligence at least be able to fold my laundry?)