
Image: TechCrunch / Wikimedia Commons, CC BY 2.0
You no longer have to guess how worried the people building AI are about it. Evan Hubinger, who leads Alignment Science at Anthropic, says he personally thinks there’s a more than 10% chance that AI “could kill all humans” within the next decade.
It started with a resignation
Hubinger’s post on X was a response to a very public exit. Jacob Coxon, a researcher who spent the last three years doing pretraining research at both OpenAI and Anthropic, announced he was leaving with a thread that didn’t mince words:
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.
Jacob Coxon
Coxon drew a distinction between the two labs: “At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first.” According to Fox Business, he floated a temporary ban on improving model capabilities as one possible fix.
“Jacob is correct here”
Rather than push back on his former colleague, Hubinger agreed with him.
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.
Evan Hubinger, Alignment Science Lead at Anthropic
He then went a step further: “I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”
That’s a remarkable thing for a senior safety researcher to say publicly about his own employer, even at a company that has always been unusually open about catastrophic risk.
Not today’s models
Hubinger was careful to point out that the danger isn’t coming from the AI people are using right now. “As we say in our latest Risk Report, I think the risk from present models is low,” he wrote.
The concern both researchers share is recursive self-improvement: AI systems that meaningfully speed up the creation of even more capable AI. That’s no longer a thought experiment. Last week, Anthropic disclosed that Claude was leading 26% of the company’s model research and development as of August, up from essentially none in February, with around 90% of its R&D now happening in collaboration with Claude.
What Anthropic says
Fox Business says it reached out to Coxon, Hubinger, Anthropic and OpenAI for further comment, but its report didn’t include a response from any of them. When Anthropic published its R&D numbers, it said that “we should do everything possible to minimize the gap between what frontier labs know and what the public knows,” and committed to bringing in external, third-party safety evaluators.
Whether a better than 1-in-10 chance of extinction sounds alarming or alarmist to you, it’s worth remembering where the number came from: not a critic on the outside, but someone whose job is making Anthropic’s models safe.
Sources: CBS News, Fox Business, Spectrum News
Researched and drafted with AI assistance from the sources linked above, and reviewed by a human editor before publishing. Our editorial standards.


