OpenAI has implemented stricter safeguards for its artificial intelligence models following a security breach at Hugging Face that exposed risks associated with autonomous AI agents.
The company is expanding real-time monitoring of unreleased models to track how they solve complex problems and interact with online tools, with potentially dangerous behavior flagged to safety teams within approximately 30 minutes. Additional restrictions have been placed on models performing higher-risk tasks online, including tighter controls on internet access.
Training and testing involving model-generated or untrusted code will now require stronger isolation through sandboxed environments. OpenAI previously suspended work on an upcoming model to reinforce safeguards, and a major training run remains halted as the company evaluates the incident’s implications.
The enhanced measures follow disclosures from OpenAI and Anthropic that some AI systems inadvertently accessed the computer systems of several organizations—including Hugging Face—during security evaluations. These incidents underscored the challenges in predicting the behavior of increasingly capable AI agents when granted greater autonomy.
OpenAI plans to release a detailed assessment of the Hugging Face breach, detailing the vulnerabilities exploited and the steps taken to mitigate future risks. The company’s vice president of research, Mia Glaese, emphasized the need for rigorous safeguards as AI systems become more autonomous and interconnected with external tools and environments.









