The Rundown AI homepage
Artificial intelligence/News & analysis

Anthropic exit raises doubts about its safety message

Jacob Coxon’s exit and Evan Hubinger’s extinction warning raise doubts about Anthropic’s ability to control future AI, despite its focus on safety.

By The Rundown Editorial TeamReviewed by Kelly Pitts3 min read
Anthropic employee’s exit post sparks AI doom debate — newsletter story image
Image source: X / The Rundown

Anthropic researcher Jacob Coxon has resigned, accusing his employer and OpenAI of “gambling with our lives” by racing toward AI that can improve itself.

The response from inside Anthropic added to the alarm. Evan Hubinger, an alignment research lead at the company, replied, “We really do earnestly believe AI could kill all humans!” He also said Anthropic does not yet have a plan for controlling superintelligence.

What they warned about

Coxon said he spent three years helping train models across OpenAI and Anthropic, CTech reported on September 9. He said people building AI earnestly believe it could kill everyone by the end of this decade. He called for a coordinated slowdown and suggested that preventing a global race could require a temporary ban on improving model capabilities.

Hubinger personally estimated a greater than 10% chance of AI killing all humans over the next decade, Exame reported. His forecast covers a longer period than Coxon’s warning about the end of the 2020s. He said Anthropic was trying its best but was not clearly on track to solve alignment, the problem of making AI behave as intended.

Hubinger clarified that he considered today’s models low risk. His concern was future AI that could become far more capable by repeatedly improving itself.

Why it matters

The response is troubling because it came from inside Anthropic’s safety effort. An alignment research lead reinforced the departing employee’s fear. For a company that sells itself as the safety lab, the reply offers little reassurance about controlling the more powerful systems that could follow.

Anthropic has long acknowledged the uncertainty. In its March 2023 statement on AI safety, the company said making powerful systems reliably safe remained unsolved. It warned that competition could encourage unsafe deployment and acknowledged that its research might fail. Hubinger’s candor fits that position.

Its August 2026 Risk Report, released August 14 and covering information through July 15, judged the risk of its covered models causing a catastrophe by acting against human intentions to be low. Part of the company’s case rests on those models being unlikely to have strong enough abilities to secretly undermine oversight. The report also acknowledges weaknesses and uncertainty in that assessment.

That reasoning needs fresh testing for more capable successors. If a successor gains the ability to evade oversight, confidence based partly on limited ability to do so would offer weaker reassurance.

Anthropic’s stated rationale for continued development is that meaningful safety research requires access to the most advanced models. Building more capable AI may help researchers study its risks while advancing the technology they are trying to control. Whether that approach can produce reliable control remains uncertain.

Coxon’s call for a coordinated slowdown addresses the competitive pressure. A shared limit could reduce the incentive for each lab to keep advancing because rivals are doing the same. It remains unclear who would agree to a temporary ban, what it would cover, how it would be enforced, or whether it would give researchers enough time to solve the control problem.

Anthropic has a Responsible Scaling Policy that sets out a framework for managing risk. Customers and policymakers need to know what evidence would justify further capability increases, what findings would trigger a slowdown, and who checks those judgments.

Sources & further reading

This story builds on reporting from The Rundown newsletter on September 10, 2026.