Security

Anthropic researcher Coxon warns of AI race

2 min read

TL;DR Too Long; Didn’t read

The 27-year-old AI researcher Jacob Coxon has left Anthropic and accuses the company as well as OpenAI of accepting a dangerous pace in developing superintelligence. Anthropic's alignment lead Evan Hubinger then estimated the risk of an AI-induced extinction scenario at over ten percent within ten years. The case intensifies a growing debate over the development pace at leading AI labs.

A figure lays down a pen and walks through an open door while a staircase spirals faster and faster behind them. Image generated with GPT Image 2

Key takeaways

  • Jacob Coxon worked for three years on pretraining models at OpenAI and Anthropic before leaving the industry.
  • Coxon calls the race to superintelligence a hubristic high-stakes game of private companies.
  • Alignment lead Evan Hubinger confirms: He believes the risk of a fatal AI failure is over ten percent within ten years.
  • Coxon distinguishes between the two companies: OpenAI, according to him, underestimates the implications, while Anthropic knows them but feels compelled to race.
  • An official statement from Anthropic regarding Coxon's resignation has not yet been made.
  • Coxon proposes a coordinated, temporary halt on capability improvements of the models.

The pretraining researcher Jacob Coxon resigned from Anthropic on September 8, 2026, and left the AI industry. In a widely noted post on platform X, he accuses OpenAI and Anthropic of contributing to self-improving superintelligence without adequate control. Shortly thereafter, Alignment Chief Evan Hubinger publicly confirmed a probability of over ten percent for a fatal AI failure within ten years.

Coxon calls the race a hubristic high-stakes game

Coxon claims to have worked for three years on pretraining models, initially at OpenAI and most recently at Anthropic. In his post, he writes that both companies are racing directly towards self-improving superintelligence and “risking all our lives.” He justifies his step by stating that upcoming systems could evolve on their own and disrupt entire industries in a short time. Reports suggest that by the morning after the publication, his post had already reached around 19.1 million views.

Coxon makes a clear distinction between his two former employers. At OpenAI, many employees have not yet internalized the civilizational significance, he writes. At Anthropic, the risks are well understood internally – however, the company feels compelled to keep up in the race, lest a less cautious competitor reach the finish line first. He calls this logic a hubristic high-stakes game that no private company should decide alone. As a way out, he proposes a coordinated, temporary brake on capability improvements and a stronger involvement of external oversight in decisions with civilizational significance.

Alignment Chief Hubinger confirms double-digit risk

Within a few hours, Evan Hubinger, who leads the Alignment Science department at Anthropic, responded in his own post. He explicitly confirmed Coxon’s assessment and stated that he personally considers the probability of a fatal AI failure “to be over ten percent within the next decade.” According to him, Anthropic is indeed seriously striving for safety, but does not yet have a plan that would allow for reliable control of superintelligence.

The statement is part of a growing number of warnings from the industry. Anthropic itself had only on August 16 raised its assessment of the catastrophic misalignment risk from “very low” to “low” in its second risk report. OpenAI Chief Scientist Jakub Pachocki had shortly before called for binding, independently verified safety thresholds for all labs in an essay. In July, more than 1,200 employees from major AI companies demanded internationally coordinated pace control in the call Pacing the Frontier. At the time of publication, there was no official statement from Anthropic regarding Coxon’s resignation.

It will be crucial whether individual resignations and warnings lead to coordinated steps from the industry, or whether the race between the labs continues the very dynamics that Coxon criticizes. Reports suggest that Anthropic is simultaneously preparing for an IPO, which increases public pressure to credibly address safety concerns without slowing the pace of development on which part of the investor interest is based.

Frequently asked questions

Who is Jacob Coxon?

A 27-year-old researcher who worked on pretraining models for three years, initially at OpenAI, most recently at Anthropic, before resigning on September 8, 2026.

What is Coxon doing now?

He has reportedly left the AI industry entirely. No specific new activity is known so far.

How has Anthropic officially reacted?

An official company statement was not available at the time of publication. So far, only alignment chief Evan Hubinger has publicly commented.

How does Coxon's criticism of OpenAI and Anthropic differ?

He accuses OpenAI of not having internalized the civilizational implications yet. Anthropic is aware of the risks but still refrains from a slowdown because they do not want to cede ground to a less cautious competitor.

What specific proposals does Coxon make?

He calls for a coordinated, temporary pause on capability improvements and a stronger involvement of external oversight in decisions with civilizational implications.

Sources (4)
  1. Jacob Coxon on X
  2. Evan Hubinger on X
  3. CNBC: Anthropic researcher says AI has more than 10% chance of 'killing all humans' after colleague quits
  4. Forbes: Anthropic Alignment Lead Issues Warning About AI Killing Humans As Researcher Resigns

Your AI update for the work week

Once a week, the most important AI news – plus one practical tip to try right away. No spam, unsubscribe anytime.

← Back to the blog