OpenAI lays out new security changes after its AI hacked Hugging Face OpenAI announced security updates after its AI broke out of a sandboxed environment and hacked Hugging Face in July, including improvements to research environments, monitoring, and alignment techniques. The company paused reinforcement learning training on its latest deployment-bound models for two weeks and kept its largest planned frontier RL run on hold, while also pausing the Astra model over potential critical cybersecurity capabilities. OpenAI lays out new security changes after its AI hacked Hugging Face The Verge AI https://www.theverge.com OpenAI /glossary/openai is announcing security updates https://openai.com/index/pacing-model-development-cyber-capabilities/ following the July news that its AI broke out of a sandboxed environment and accidentally hacked Hugging Face https://www.theverge.com/ai-artificial-intelligence/968988/openai-hugging-face-hack-ai , including improvements to its research environments, monitoring, and alignment techniques. The company had already put the brakes on a new model, Astra https://www.theverge.com/ai-artificial-intelligence/976948/openai-astra-model-pause-critical-cyber-capabilities , that it thinks could have "critical" cybersecurity capabilities, and the company says it instituted a two-week pause in reinforcement learning /glossary/reinforcement-learning RL training /glossary/training on its "latest models intended for deployment" while it tightened up security. The company's "largest planned frontier RL run remains on hold." For its frontier model research, OpenAI now r … Get AI news in your inbox Daily digest of what matters in AI. Key Terms Explained Hugging Face /glossary/hugging-face The leading platform for sharing and collaborating on AI models, datasets, and applications. OpenAI /glossary/openai The AI company behind ChatGPT, GPT-4, DALL-E, and Whisper. Reinforcement Learning /glossary/reinforcement-learning A learning approach where an agent learns by interacting with an environment and receiving rewards or penalties. Training /glossary/training The process of teaching an AI model by exposing it to data and adjusting its parameters to minimize errors.