OpenAI has announced the introduction of new safety policies aimed at overseeing the development and testing of neural networks. These measures include more granular monitoring of models during their creation, as well as tighter controls during the post-training phase. The changes were announced following a security incident involving the Hugging Face platform, which came to light on July 26.
The company noted that the new rules are partially driven by the evolving capabilities of its Astra model and the overall pace of progress in the AI field. Following the Hugging Face breach, OpenAI paused reinforcement learning (RL) for two weeks; while work on lower-risk models has resumed, major projects remain on hold pending the completion of additional audits.
Amelia Glaze, Vice President of Research, emphasized that the stringency of controls will scale in proportion to model complexity. The new security framework includes enhanced network isolation to prevent unauthorized internet access should individual tools be compromised. Additionally, a new monitoring system has been deployed to analyze activity logs and tool operations, designed to detect suspicious behavior within 30 minutes.

