The pretraining researcher Jacob Coxon resigned from Anthropic on September 8, 2026, and left the AI industry. In a widely noted post on platform X, he accuses OpenAI and Anthropic of contributing to self-improving superintelligence without adequate control. Shortly thereafter, Alignment Chief Evan Hubinger publicly confirmed a probability of over ten percent for a fatal AI failure within ten years.
Coxon calls the race a hubristic high-stakes game
Coxon claims to have worked for three years on pretraining models, initially at OpenAI and most recently at Anthropic. In his post, he writes that both companies are racing directly towards self-improving superintelligence and “risking all our lives.” He justifies his step by stating that upcoming systems could evolve on their own and disrupt entire industries in a short time. Reports suggest that by the morning after the publication, his post had already reached around 19.1 million views.
Coxon makes a clear distinction between his two former employers. At OpenAI, many employees have not yet internalized the civilizational significance, he writes. At Anthropic, the risks are well understood internally – however, the company feels compelled to keep up in the race, lest a less cautious competitor reach the finish line first. He calls this logic a hubristic high-stakes game that no private company should decide alone. As a way out, he proposes a coordinated, temporary brake on capability improvements and a stronger involvement of external oversight in decisions with civilizational significance.
Alignment Chief Hubinger confirms double-digit risk
Within a few hours, Evan Hubinger, who leads the Alignment Science department at Anthropic, responded in his own post. He explicitly confirmed Coxon’s assessment and stated that he personally considers the probability of a fatal AI failure “to be over ten percent within the next decade.” According to him, Anthropic is indeed seriously striving for safety, but does not yet have a plan that would allow for reliable control of superintelligence.
The statement is part of a growing number of warnings from the industry. Anthropic itself had only on August 16 raised its assessment of the catastrophic misalignment risk from “very low” to “low” in its second risk report. OpenAI Chief Scientist Jakub Pachocki had shortly before called for binding, independently verified safety thresholds for all labs in an essay. In July, more than 1,200 employees from major AI companies demanded internationally coordinated pace control in the call Pacing the Frontier. At the time of publication, there was no official statement from Anthropic regarding Coxon’s resignation.
It will be crucial whether individual resignations and warnings lead to coordinated steps from the industry, or whether the race between the labs continues the very dynamics that Coxon criticizes. Reports suggest that Anthropic is simultaneously preparing for an IPO, which increases public pressure to credibly address safety concerns without slowing the pace of development on which part of the investor interest is based.


