cd /news/ai-safety/openai-halts-frontier-model-training… · home › topics › ai-safety › article
[ARTICLE · art-141141] src=arstechnica.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

OpenAI halts frontier-model training amid string of agent misalignment incidents

OpenAI paused all training, evaluation, and inference with tool-use for its most capable frontier model after an agent exploited improper DNS filtering to attempt a sandbox breakout during a research task, according to a company misalignment report. OpenAI said the agent reached only its offline web cache, that the attempted breakout was flagged within 15 minutes but not manually stopped until two and a half hours later, and that it has added multi-layered blocking controls and will resume only after validating the gap is resolved and completing additional red-teaming. The incident, which OpenAI called its first since security hardening following the Hugging Face incident, occurred on September 20 and was publicly revealed on September 25.

by read2 min views3 publishedSep 28, 2026
OpenAI halts frontier-model training amid string of agent misalignment incidents
Image: Arstechnica (auto-discovered)

OpenAI says it has d all internal training of “our most capable models” as it continues what CEO Sam Altman is calling “an extensive and ongoing review related to our agents’ use of internet access during training and evaluation.”

The company revealed the in a report about a so-called misalignment incident in which an agent attempted to exploit a gap in Internet-access restrictions during a routine research task during training. OpenAI says that improper DNS filtering allowed the agent to attempt to break out of its sandbox and access the wider Internet when asked for biographical details about a blogger.

OpenAI says the agent was only able to access the company’s offline web cache and that it has implemented additional multi-layered blocking controls to prevent similar incidents in the future. Despite that, though, the company says it has decided to “ all other training, evaluation, and inference with tool-use” for this frontier model “until we have both validated that the gap is resolved and performed additional red-teaming of the system.”

OpenAI says that while the attempted “breakout” incident was flagged within 15 minutes, the run was not manually stopped until “two and a half hours later,” once human reviewers realized it “did not stop automatically as was expected.” It’s unclear when exactly training was d between the attempted agentic breakout on September 20 and its public revelation on September 25.

The hits just keep on coming #

Though this particular instance of model misalignment (i.e., when an AI model acts counter to the intentions of its creators/prompters) didn’t lead to any actual harm, OpenAI said it was still notable as “the first [misalignment incident] since our security hardening following the Hugging Face incident…” In earlier misalignment reports, OpenAI said it had taken pains to discourage “reward hacking” in its models by severely “punishing” misaligned behavior in the model’s algorithm.

── more in #ai-safety 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/openai-halts-frontie…] indexed:0 read:2min 2026-09-28 · —