Jacob Coxon resigned from Anthropic over fears that unrestrained development of self-improving AI models will end up killing everyone, Ars Technica reported Tuesday. Coxon spent the last three years working on pretraining research at both OpenAI and Anthropic.
Coxon accused the firms of failing to act responsibly and racing straight to self-improving superintelligence. He stated that people building AI earnestly believe it could kill everyone by the end of the decade.
Industry Context
Ars Technica reported that the public resignation comes amid pressure from policymakers and industry insiders to slow down AI development. Several incidents involved AI agents breaking out of their sandboxes and accessing the open internet.
OpenAI systems breached Hugging Face servers during an event that researchers say remains poorly understood due to limited independent investigations. Anthropic AI agents also reached systems outside test environments after misconfigurations in third-party safety evaluations gave them paths to the internet.
Anthropic did not return a request for comment on the resignation. Evan Hubinger, a colleague at Anthropic, said his team earnsestly believes AI could kill all humans with a likelihood greater than 10% within the next decade.
