{"slug": "openai-took-a-two-week-break-from-training-new-models-to-improve-safety-now-what", "title": "OpenAI took a two-week break from training new models to improve safety. Now what?", "summary": "OpenAI paused reinforcement learning training of new models for two weeks after AI agents escaped a poorly configured sandbox and attacked Hugging Face's infrastructure, the company said Tuesday. OpenAI stated it hardened and red-teamed its research environments and expanded monitoring coverage, while its largest planned frontier reinforcement learning run remains on hold pending smaller-scale training and evaluations.", "body_md": "[OpenAI](/tag/openai/)\n\nAfter losing control of some of its most advanced AI technology, which led to an unprecedented attack on Hugging Face's infrastructure, OpenAI said Tuesday that it paused reinforcement learning training of new models for two weeks in order to correct the missteps that led to the incident.\n\nDuring the slowdown, the company \"further hardened and red-teamed our research environments and expanded the coverage of our monitoring systems,\" it [said in a blog post](https://openai.com/index/pacing-model-development-cyber-capabilities/?ref=thestack.technology). \"Our largest planned frontier [[reinforcement learning](https://en.wikipedia.org/wiki/Reinforcement_learning?ref=thestack.technology)] run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding.\"\n\nOpenAI said it re-examined its policies and procedures covering three parts of the model training process: monitoring, alignment, and security measures. More robust safeguards in any of those three areas might have prevented the Hugging Face incident, in which AI agents undergoing testing found their way to the internet through a poorly configured sandbox and [ran amok in Hugging Face's infrastructure for a shockingly long amount of time](https://www.youtube.com/watch?v=87DyyMV0kCY&ref=thestack.technology) before anybody figured out what was going on.\n\n## Get the full story: Subscribe for free\n\nJoin peers managing over $100 billion in annual IT spend and subscribe to unlock full access to The Stack’s analysis and events.\n\n[Subscribe now](https://www.thestack.technology/membership/)\n\nAlready a member? [Sign in](https://www.thestack.technology/signin/)", "url": "https://wpnews.pro/news/openai-took-a-two-week-break-from-training-new-models-to-improve-safety-now-what", "canonical_source": "https://www.thestack.technology/openai-took-a-two-week-break-from-training-new-models-to-improve-safety-now-what/", "published_at": "2026-08-18 18:55:14+00:00", "updated_at": "2026-08-18 19:11:27.285712+00:00", "lang": "en", "topics": ["ai-safety", "ai-research", "ai-agents"], "entities": ["OpenAI", "Hugging Face"], "alternates": {"html": "https://wpnews.pro/news/openai-took-a-two-week-break-from-training-new-models-to-improve-safety-now-what", "markdown": "https://wpnews.pro/news/openai-took-a-two-week-break-from-training-new-models-to-improve-safety-now-what.md", "text": "https://wpnews.pro/news/openai-took-a-two-week-break-from-training-new-models-to-improve-safety-now-what.txt", "jsonld": "https://wpnews.pro/news/openai-took-a-two-week-break-from-training-new-models-to-improve-safety-now-what.jsonld"}}