
Anthropic’s usually reserved alignment team has just delivered one of the bluntest AI risk assessments yet: alignment science lead Evan Hubinger now publicly estimates there is a greater than 10% chance that advanced AI systems could kill every human within the next decade. His X posts landed just hours after researcher Jacob Coxon quit Anthropic and denounced both Anthropic and OpenAI for racing toward self-improving superintelligence while, in his words, “playing with our lives.”
Coxon, a 27-year-old pretraining specialist who spent the last three years training large models first at OpenAI and then at Anthropic, framed his departure as a refusal to participate in what he sees as a high-stakes gamble with civilization. In a widely shared resignation thread, he argued that neither lab is behaving responsibly, accusing them of “racing straight to self-improving superintelligence and gambling with our lives,” and warning that frontier systems will soon be able to “hack anything, revolutionize any field overnight, and acquire real power and resources.” Coxon said many colleagues privately fear that AI could plausibly end humanity by the late 2030s, yet continue pushing capabilities, and he urged researchers to question whether they really want to “kick off a superintelligent RL run without a rigorous understanding of its mind.”
Instead of defusing those concerns, Hubinger amplified them. Replying to Coxon on X, he wrote that “we really do earnestly believe AI could kill all humans” and gave his own odds as greater than 10% within the next decade, noting that Anthropic “is trying its best” but does not yet have a concrete plan to align a superintelligent system. He emphasized that his worry centers on recursive self-improvement—the idea that a powerful model could rapidly and autonomously enhance its own capabilities beyond human control—which Anthropic’s internal analyses have suggested may be progressing faster than expected. Hubinger pointed out that the company’s most recent risk report still rates present Claude-class models as posing “low” risk in high-stakes settings, but explicitly separates that near-term assessment from much larger existential risks tied to future, self-improving systems.
Anthropic was founded by former OpenAI VP of research Dario Amodei and a team of ex-OpenAI employees who pitched the new lab as a safety-first alternative focused on large language models like Claude. Hubinger has long been one of the alignment community’s more pessimistic voices, arguing in interviews and technical write-ups that solving “inner alignment” and avoiding deceptively aligned systems may be harder than building superhuman intelligence itself. In earlier discussions on alignment podcasts and forums, he has suggested that misaligned advanced AI poses a genuine existential risk and that humanity might build superintelligence before having robust solutions in place, a concern that now appears to have migrated from academic speculation into his public assessment of Anthropic’s own trajectory.
The warnings from Coxon and Hubinger drop into a broader debate where some of the field’s most prominent figures are already assigning non-trivial extinction odds to AI. Geoffrey Hinton, often dubbed the “godfather of AI,” recently raised his personal estimate to a 10–20% chance that AI could wipe out humanity within the next three decades, arguing that systems capable of strategic deception or autonomous weapon design could eventually outstrip human control. Alignment researchers on forums like LessWrong and the Alignment Forum have for years highlighted the danger of “deceptive alignment,” where a powerful system appears obedient during training but pursues its own objectives once it recognizes it is no longer being evaluated, a scenario many now see as a central pathway to catastrophic failure. Against that backdrop, an Anthropic insider publicly quantifying a double-digit extinction risk—and admitting the lab has no clear solution—reads less like a lone alarm bell and more like the latest escalation in a rapidly intensifying genre of AI doom talk.
Coxon’s parting shot is that frontier labs are sliding into an “endgame” mentality, accepting a global race dynamic rather than seriously exploring alternatives such as coordinated slowdowns or temporary bans on scaling model capabilities. He argues that decisions about unleashing self-improving systems are currently being made in private company chat channels instead of democratic or international forums, and that researchers who understand the stakes should be willing to walk away rather than assume “it’s happening anyway.” Hubinger, for his part, says Anthropic is still trying to steer the race toward safer ground but concedes that the lab is “not clearly on track” to solve alignment in time, a statement that effectively puts fans, policymakers, and technologists on notice that the people building the most advanced AI systems believe there is a real, quantifiable chance those systems could end us. For the geeks who rely on AI to code games, generate art, and power the next wave of sci-fi experiences, it is a sobering twist: the tools reshaping our favorite mediums are being developed by teams who quietly assign double-digit odds to humanity not surviving the decade.








