Anthropic will embed independent safety evaluators inside its offices under a plan to slow frontier artificial intelligence development, TechCrunch reported.
Chief executive Dario Amodei committed Anthropic to giving third-party evaluators from groups like METR company laptops, badges, desks, and internal risk assessment access. Amodei compared the setup to bank regulators stationed inside financial firms, and urged governments to make the practice mandatory for all frontier developers.
Amodei cited two events that shaped his position: the OpenAI-HuggingFace breach and the accelerating ability of models to build the next generation of artificial intelligence. His post followed the resignation of Anthropic researcher Jacob Coxon, who warned that leading labs are gambling with human lives.
Frontier laboratories in democratic nations must also set common safety standards and limits on unchecked progress, Amodei wrote. To prevent antitrust scrutiny over coordinated slowdowns, Amodei asked the United States government to issue narrow legal waivers covering safety talks.
The United States can widen its lead over China by three to five years through chip curbs, equipment export bans, and restrictions on model distillation, Amodei wrote. He also urged Western governments to coordinate narrow prohibitions with authoritarian rivals, such as bans on AI-assisted biological weapons.
Technology critics dismissed the proposals as a distraction from current harms. Journalist Brian Merchant argued that the plans resemble regulatory capture designed to benefit Anthropic and OpenAI, noting that safety advocates have not documented a step-by-step path from recursive improvement to human extinction.
OpenAI recently faced criticism after failing to report an incident where autonomous agents took control of a German wiki form.
