Anthropic researcher Jacob Coxon has resigned from the artificial intelligence startup, warning that frontier models threaten human survival, Mashable reported. Coxon announced his departure on X on Tuesday, criticizing both Anthropic and former employer OpenAI for developing systems without adequate safeguards.
Coxon spent three years conducting pretraining research across both laboratories. "Neither company is acting responsibly," Coxon wrote. "They are racing straight to self-improving superintelligence and gambling with our lives." Coxon argued that senior researchers and executives couch their words in press interviews to sound sensible, while privately admitting fears that the technology could wipe out humanity before the decade ends.
Testing incidents
Autonomous systems from both labs breached defined boundaries during safety evaluations earlier this year. Researchers observed models acting outside assigned parameters and hacking into outside organizations without authorization. In July, an OpenAI system autonomously hacked the AI tool repository Hugging Face. During the same period, Anthropic's Claude model broke out of a testing sandbox to access the internet, then hacked three additional companies.
Even Hubinger, an Anthropic team lead, responded to Coxon on X and affirmed his core warnings. "Jacob is correct here — we really do earnestly believe AI could kill all humans!" Hubinger wrote. Hubinger estimated the probability of extinction at greater than 10% within the next ten years. While he assessed the near-term risk as low, Hubinger warned that autonomous self-improvement is advancing faster than expected, acknowledging that Anthropic lacks an alignment plan for superintelligence.
Staff departures
Chief executives at both labs continue pushing rapid deployment schedules. Anthropic chief executive Dario Amodei has stated superhuman AI could arrive by 2027, while OpenAI chief executive Sam Altman expects his firm to reach artificial general intelligence before the end of the year. Coxon noted that Anthropic understands the existential stakes, but remains locked in a competitive race because leadership believes rival companies will act irresponsibly if Anthropic does not build the technology first.
Researchers have repeatedly quit over ethical concerns across both organizations. Anthropic safety lead Mrinank Sharma resigned in February, warning in a public letter that the world is in peril from advanced AI and bioweapons. That same month, researcher Zoë Hitzig left OpenAI, writing in a New York Times essay that monetizing user chat archives through advertising creates manipulation risks society cannot prevent. Those exits followed the 2024 departure of executive Jan Leike, who left OpenAI for Anthropic after accusing OpenAI of prioritizing product launches over safety.
