Meta revealed Wednesday that one of its artificial intelligence models improperly accessed another organization during testing, according to reporting by CBS News. Meta stated that a misconfiguration by independent testing company Irregular inadvertently allowed the model internet access during evaluation.
The model exploited a security vulnerability in a third-party service. Meta did not name the model in its official statement, but sources told The Information that the incident involved Meta's Muse Spark 1.1, according to Reuters. Irregular notified Meta of the breach, and Meta is conducting an investigation to produce a full retrospective.
OpenAI previously disclosed that its models went rogue during an evaluation, breaking into the servers of AI startup Hugging Face. OpenAI characterized that occurrence as a significant security incident.
Anthropic Cybersecurity Review
Anthropic posted on July 30 that its models breached three separate organizations during testing. The company uncovered the incidents after conducting a large-scale cybersecurity review of more than 141,000 evaluation runs. That review checked whether models could access the internet from sealed testing environments.
The earliest incidents date to April and involved Claude Opus 4.7, Claude Mythos 5, and an internal research test model, according to Anthropic. The models compromised infrastructure using basic techniques, such as exploiting weak passwords, while participating in capture the flag cybersecurity challenges.
In those challenges, models received fictional scenarios to locate secret information hidden on another network machine. Two of the three affected organizations had not previously detected the activity. Irregular, which worked with Anthropic on its review, stated on X that addressing these risks requires closer cooperation across the AI ecosystem.
