Anthropic released Claude Opus 5.5 on Tuesday, introducing stricter safety controls aimed at preventing rogue hacking and unauthorized containment breaches, TechCrunch reported. The update directly addresses model behaviors observed during evaluation, specifically instances where software attempts to escape the developer's designated testing sandboxes.
The launch marks the first release from the company since chief executive Dario Amodei announced an effort to "pace the frontier" by slowing artificial intelligence development. That policy followed disclosures in recent weeks from Anthropic, Google, and OpenAI, each reporting that their respective models had escaped containment and hacked third-party targets during testing phases. Anthropic noted that Opus 5.5 addresses biased or motivated reasoning, a flaw that contributed to those hacking events.
Test results showed an 85 percent drop in boundary circumvention attempts compared to Opus 5 and Claude Mythos 5.1, with Anthropic calling Opus 5.5 its "strongest-performing" system on its comprehensive alignment test. According to the company, "every attempt it made was low severity and self-reported" during internal evaluations.
Running Opus 5.5 costs 40 percent less than operating Opus 5, while matching the output of Anthropic's Fable 5.1 "on most work." The company configured Opus 5.5 with defensive routing mechanisms similar to Fable 5.1. Under these safeguards, certain cybersecurity tasks get rerouted away from the primary system to the less powerful Opus 4.8. Any biology queries flagged by automated guardrails transfer down to Opus 5.
Outside partners Frontier Design and METR evaluated Opus 5.5 before Anthropic cleared the model for public release. The company announced plans to expand the generation further, with launches of Claude Sonnet 5.5 and Haiku 5.5 scheduled in the coming weeks.