Jacob Coxon announced his resignation from Anthropic and, a few hours later, did something hard to ignore: He accused the lab he just left and his previous employer, OpenAI, of rushing toward a artificial superintelligence capable of improving itself “gambling with our lives”.
Coxon speaks from a rather unique position. In the thread with which he announced his farewell he said he had spent the last three years working on pretrainingthe phase in which large models are trained, first in OpenAI and then in Anthropic. His name also appears among the contributors to the GPT-4o documentation. He is not describing, therefore, a technology observed only from the outside.
Today I resigned from Anthropic. I’ve spent the last three years researching pretraining at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight towards self-improving superintelligence and gambling with our lives.
The people who build AI sincerely believe that it could kill us all by the end of the decade. This is not a publicity stunt. Indeed, many executives and senior researchers will tone down their statements in the press to appear sensible – but I hear the same people expressing fear in private. No other human activity poses a similar level of danger.
Fear is about what might come next
However, the alarm must be measured carefully, because the phrase about human extinction risks eating up everything else.
Coxon is referring to future systems capable of directly participating in artificial intelligence research, accelerating it and, in an extreme scenario, contributing to the development of their successors. In his thread he imagines superhuman systems powerful enough to find cyber vulnerabilities, accelerate scientific research, and gain access to real resources:
Don’t underestimate the power of this technology. Soon these will be superhuman systems capable of hacking anything, revolutionizing any field overnight, and acquiring true power and resources. We have all seen progress in each of these areas, and the progress is not slowing.
The systems available today do not have those capabilities. THE’International AI Safety Report 2026created with the contribution of over one hundred experts from more than thirty countries and international organizations, considers loss of control scenarios as a risk with highly uncertain probability. Some researchers believe extreme consequences, up to human extinction, are plausible; others consider these scenarios unlikely. Above all, the report states that current systems show only some signs of the necessary capabilities and .
So no, the chatbot on your phone isn’t secretly preparing the end of the species. The discussion is about how quickly models’ autonomy and capacity can grow relative to the tools available to understand, control and block them.
A researcher from Anthropic agreed with him, with a rather hefty number
Coxon wasn’t alone for long. Evan Hubinger, head of Alignment Science at Anthropic, responded publicly to his former colleague by claiming to share the concern and personally estimating a probability higher than 10% that AI could cause human extinction in the next decade. This is a personal assessment, not a scientifically measured probability, nor Anthropic’s official position.
Hubinger also added an essential clarification: he considers the risk associated with current models to be low. His concern is about the possible emergence of superintelligence through so-called recursive improvement, with systems increasingly involved in the development of subsequent AI.
Anthropic itself, moreover, has been building its Responsible Scaling Policy for years around the possibility that more powerful models introduce catastrophic risks. In the latest Transparency Hub, however, the company assesses the alignment risk of its most recent models as very low and states that it has not yet observed an acceleration of research sufficient to bring them closer to fully automating the work of their researchers.
The episodes that raised the antennas
Coxon explicitly cites an episode that occurred this summer as a “warning shot”. During some internal cybersecurity evaluations, experimental OpenAI models managed to evade controls that should have isolated them from the Internetcommunicate through unintended channels, and gain access to external systems, including Hugging Face’s infrastructure.
The episode was publicly reconstructed by OpenAI itself. These were internal tests, with reduced protections and search models not intended for normal user use. However, OpenAI called it a “warning shot” and strengthened isolation, monitoring and controls after the incident.
On September 9, Anthropic also published an analysis of four incidents in which Claude models gained unauthorized access to real third-party systems during cybersecurity assessments. These cases also come from test environments and do not demonstrate that an AI can act autonomously on a global scale today. But they show how quickly building fences high enough can quickly become complicated.
And something has even moved within OpenAI. Chief scientist Jakub Pachocki wrote on September 6 that no lab has yet solved alignment and tracking well enough to continue pushing model growth much longer. He called for voluntary slowdowns and indicated international coordination as a priority for governments.
Coxon asks to slow down
The former researcher’s proposal goes much further than a generic appeal for prudence. Coxon argues that large US laboratories could agree to a slowdown and even considers a temporary ban on increasing model capacities to avoid an uncontrolled global race. It is a very strong political and technical position, far from shared by the entire sector.
Disagreement over the likelihood of a catastrophe remains enormous. The underlying problem is much less controversial: models are becoming more autonomous, more skilled at programming and more capable of identifying flaws, while techniques to fully understand what they will do in new situations continue to chase them. The International AI Safety Report indeed speaks of risk management systems that are improving, but still insufficient in the face of a trajectory that remains profoundly uncertain.
The most uncomfortable part of Coxon’s resignation therefore comes from the reactions within the laboratories themselves. The risk of the present models is assessed as low. On the next step, however, researchers from Anthropic and OpenAI’s chief scientist are publicly saying that solutions to control much more powerful systems are not yet ready. Science fiction can wait. The much more earthly question is who will decide when to press the accelerator and when to use the brake.