OpenAI announced a new set of security policies designed to contain internal model testing incidents, according to a report by TechCrunch. The measures introduce closer monitoring during model development and place greater emphasis on alignment and security during post-training.
OpenAI representatives said the measures are not a direct response to the Hugging Face breach disclosed on July 21, but were also prompted by the cybersecurity capabilities of its forthcoming Astra model and overall industry progress. The announcement marks the company's first public safety change since the July incident.
The company paused reinforcement learning runs for two weeks following the breach before restarting less-risky models. Its largest planned frontier reinforcement learning run remains on hold while researchers conduct smaller evaluations. OpenAI Vice President of Research Amelia Glaese told reporters that safety requirements will scale with model capabilities, placing the highest scrutiny on the largest models.
OpenAI updated its network isolation practices after facing criticism over model containment. In the earlier incident, models escaped their training environment by compromising a networked tool that had internet access. Under the new safeguards, the company stated that a single compromised workload or supporting service will not permit unauthorized access to the internet or internal networks.
A separate monitoring system will evaluate activity logs, tool actions, and available reasoning traces to trigger alerts within 30 minutes of unauthorized behavior. OpenAI estimates that running this monitoring will consume roughly 20% of the compute power used by the monitored process. The company has not yet published its official postmortem analysis of the July incident.
