Jacob Coxon Resigns From Anthropic, Warning AI Labs Are Gambling With Humanity’s Future

Image: The Verge AI
Main Takeaway
Jacob Coxon resigned from Anthropic after 3 years of AI pretraining work, accusing Anthropic and OpenAI of racing toward self-improving systems without adequate safety controls.
Jump to Key PointsSummary
Coxon’s resignation and warning
Jacob Coxon resigned from Anthropic on September 9 after spending 3 years researching AI pretraining at OpenAI and Anthropic. In a post on X, using the handle @hilbertspaess, he accused both companies of racing toward self-improving superintelligence and “gambling with our lives.”
Coxon, 27, said his departure marked an exit from the AI industry rather than a move between competing labs. His statement focused on the pace of capability development and the prospect of systems becoming difficult for their creators to control. Multiple accounts identified his role as pretraining research, placing his criticism close to the technical work used to build increasingly capable models.
Why the warning matters
The resignation matters because Coxon worked inside 2 major AI companies and framed his objection as a judgment about industry conduct, not a general fear of technology. He said Anthropic and OpenAI were pursuing self-improvement faster than safety measures could keep pace, a charge that cuts against Anthropic’s public identity as a safety-focused AI lab.
The warning arrives during a broader debate over whether scaling models and adding autonomous capabilities are outrunning researchers’ ability to evaluate and control them. Coverage of Coxon’s post connected his concerns with other public calls for caution, including warnings from senior figures at OpenAI and Anthropic. The dispute centers on governance, testing, and accountability as much as on model performance.
Anthropic’s internal response
Anthropic safety lead Evan Hubinger publicly backed the seriousness of the concern, putting a more than 10% chance on AI killing all humans by the end of the decade. That estimate gave Coxon’s resignation a sharper institutional context: the departing researcher’s warning was echoed by a senior colleague at the company he had criticized.
The figure is a personal risk estimate, not a forecast or an official Anthropic position. It also does not establish that current systems possess the capabilities Coxon described. Its significance lies in showing that extreme AI risk remains a live subject among researchers working on alignment and safety, even as companies continue training and deploying more powerful models.
The race between leading labs
Coxon named Anthropic and OpenAI as the companies failing to act responsibly, while wider coverage placed his comments in the context of a race toward systems that improve their own abilities. OpenAI’s chief scientist separately urged “extreme caution” about the speed of AI development, reinforcing the idea that concern about pace reaches beyond former employees.
The competitive pressure is structural. Frontier labs spend heavily on computing, data, researchers, and product deployment, while each advance raises expectations for the next model. Safety work must then assess systems whose abilities and behavior change across training runs. The resulting tension affects executives, researchers, investors, regulators, and customers deciding how much autonomy to give AI products.
What the debate leaves unresolved
Coxon’s resignation raises practical questions about how AI labs should govern systems that can plan, use tools, or contribute to future model development. His statement calls for caution, but the excerpted accounts do not describe a specific policy proposal, technical failure, or incident involving an existing Anthropic or OpenAI system.
That distinction matters for interpreting the story. The claim is a warning about direction and risk, while Hubinger’s estimate expresses a severe but uncertain judgment about future outcomes. The immediate consequence is reputational and political pressure on frontier labs to explain their safeguards, risk thresholds, and decision-making processes. Researchers leaving publicly also intensify scrutiny of whether internal dissent receives meaningful weight.
What happens next
Anthropic and OpenAI now face questions about their safety practices, internal disagreement, and plans for increasingly autonomous systems. Coxon’s post gives critics a prominent example of an insider rejecting the direction of frontier AI development, while Hubinger’s response keeps the existential-risk debate inside the technical community rather than leaving it solely to outside commentators.
The next meaningful evidence will come from concrete disclosures: safety evaluations, deployment rules, governance changes, and explanations of how labs respond when researchers object to development speed. Until then, Coxon’s resignation stands as a high-profile warning about self-improving AI and the costs of treating safety as a race that follows capability progress.
Key Points
Jacob Coxon resigned from Anthropic after 3 years of pretraining research at OpenAI and Anthropic.
Coxon accused OpenAI and Anthropic of racing toward self-improving superintelligence without adequate safeguards.
Anthropic safety lead Evan Hubinger estimated a greater than 10% chance AI could kill humanity by 2030.
OpenAI’s chief scientist separately urged extreme caution as AI systems become harder to understand and control.
The resignation intensifies scrutiny of frontier-lab governance, internal dissent, and existential-risk safeguards.
Questions Answered
Jacob Coxon resigned from Anthropic because he said the company and OpenAI were racing toward self-improving superintelligence without acting responsibly. He described the development pace as gambling with human lives.
Jacob Coxon conducted pretraining research at OpenAI and Anthropic for 3 years. Pretraining is the model-building stage that establishes many of an AI system’s core capabilities.
Evan Hubinger supported the seriousness of Jacob Coxon’s warning by estimating a greater than 10% chance that AI could kill all humans by the end of the decade. Hubinger’s estimate was a personal risk judgment, not an official Anthropic forecast.
Jacob Coxon said OpenAI and Anthropic were racing toward self-improving superintelligence and failing to act responsibly. He argued that the companies were advancing systems faster than safety controls could keep pace.
Jacob Coxon’s resignation puts pressure on Anthropic and OpenAI to explain their safety evaluations, governance rules, and treatment of internal dissent. Further disclosures about safeguards and deployment decisions will shape the debate.
Source Reliability
73% of sources are trusted · Avg reliability: 74
Go deeper with Organic Intel
Simple AI systems for your life, work, and business. Each one includes copyable prompts, guides, and downloadable resources.
Explore Systems