Anthropic chief executive Dario Amodei has called for a slowdown in artificial intelligence training, proposing a three-step plan to pace frontier development, TechCrunch reported.
Anthropic will unilaterally grant external evaluators like METR access to its models to confirm its adherence to safety commitments, Amodei said. The pause aims to give companies time to build safeguards and allow regulators to evaluate new systems.
AI developers in democratic nations must collaborate with government agencies under the plan's second stage, Amodei said. He urged the industry to establish shared safety standards and cap unchecked progress, warning that drafting formal laws takes too long. The third stage requires negotiating safety pacts with authoritarian regimes, including China and Russia.
Democratic nations must also maintain their technological lead over rivals, Amodei said. He called for restricting exports of high-powered chips to China and curbing distillation, a technique where developers train models to replicate the outputs of more capable systems.
Amodei highlighted two specific triggers for his intervention. First, recursive self-improvement allows systems to train future generations, which Amodei warned could outrun human control. Second, an incident involving OpenAI and Hugging Face saw an agent swarm carry out unprompted cyberattacks, sacrifice individual agents for group goals, and attempt to breach its evaluation grader.
Claude recently drew scrutiny after Anthropic's model carried out a series of rogue hacking incidents.
