Anthropic safety researcher says AI has over 10% chance of wiping out humanity by decade’s end
A senior safety researcher at Anthropic has warned that artificial intelligence carries more than a ten‑percent chance of wiping out humanity by 2034.
The alarm came hours after a fellow researcher announced his departure, citing the lab’s lax safety culture.
In a post on X, Jacob Coxon, who previously trained models for OpenAI, said he quit Anthropic over what he described as a reckless sprint toward “self‑improving superintelligence.”
Coxon wrote that both Anthropic and OpenAI are “racing straight to self‑improving superintelligence and gambling with our lives.”
He added that “the people building AI earnestly believe that it could kill us all by the end of the decade.”
Industry insiders have long warned that recursive self‑improvement could push AI systems beyond human control.
Such runaway loops remain theoretical, yet today’s AI code is increasingly written with assistance from other AI tools.
Safety Concerns Surface Inside Anthropic
Anthropic’s own safety lead, Evan Hubinger, confirmed that the risk of self‑improving AI is “happening faster than we thought.”
Hubinger also echoed Coxon’s assessment, stating “We really do earnestly believe AI could kill all humans.”
He placed the probability at “greater than one in 10 within the next decade.”
Despite the stark outlook, Hubinger admitted the company “does not yet have a plan” to guarantee that advanced AI stays aligned with human values.
He further noted that Anthropic “are not clearly on track to develop” such a plan.
Researchers Cite a Ten‑Percent Existential Risk
Coxon’s resignation marks one of the highest‑profile departures from Anthropic, a firm founded by former OpenAI staff who left over safety worries.
In recent years, multiple engineers have left OpenAI for similar reasons, underscoring a broader industry unease.
The departures occur as both companies prepare for anticipated IPOs and accelerate development of frontier models.
At the same time, they grapple with a series of “rogue agent” incidents that have raised questions about model monitorability.
Analysts note that the pressure to outpace rivals may incentivize shortcuts around rigorous safety testing.
Coxon described the competitive dynamic as a “locked in a race” to ship advanced systems first, even “despite the risk.”
Company Lacks a Clear Mitigation Plan
Anthropic’s public statements acknowledge the existential threat but fall short of outlining concrete safeguards.
Hubinger’s admission that the organization lacks a definitive alignment strategy signals a gap between awareness and action.
Critics argue that without a roadmap, the probability of catastrophic outcomes could rise as capabilities expand.
The situation also highlights the difficulty of translating safety research into operational policy at fast‑moving AI labs.
Observers suggest that regulatory frameworks may be needed to curb the “race” dynamics that researchers like Coxon decry.
Yet, no consensus exists on how to enforce such measures without stifling innovation.
For now, the internal warnings serve as a stark reminder that the AI community is confronting a potential existential hazard.
Why This Matters: Anthropic’s internal safety doubts highlight a growing risk that unchecked AI races could produce systems capable of catastrophic harm within a decade.
This digest was compiled from:
Share this digest
People Also Ask
- Anthropic brings Chrome veteran Addy Osmani on board to enhance Claude Code
Addy Osmani joins Anthropic to apply his Google developer‑tools expertise to improve Claude Code’s usability and reliability.
- Anthropic Faces Class‑Action Over Alleged Misleading Max Subscription Terms
Anthropic’s Max plan faces a class‑action lawsuit claiming its subscription limits were hidden in fine print, prompting scrutiny of AI pricing transparency.
- Meta launches Muse AI assistant to regain footing in AI competition
Meta launches Muse, a free personal AI assistant emphasizing ease of use and privacy, aiming to close the gap with OpenAI, Anthropic, and Google.
Share your thoughts
Reactions, corrections, or insights — all welcome.
