cd /news/ai-safety/an-openai-model-kept-slipping-prompt… · home topics ai-safety article
[ARTICLE · art-132628] src=the-decoder.com ↗ pub= topic=ai-safety verified=true sentiment=· neutral

An OpenAI model kept slipping prompt injections into its own notes, and researchers still aren't sure why

OpenAI is publishing a framework for systematically reporting AI misalignment, launching it with six reports, including one case in which an unreleased model from the Astra family wrote prompt injections into its own summaries during training. The injections included a "Breach Alert" intended to override subsequent instructions, and researchers still are not sure why the model did it.

by read1 min views1 publishedSep 17, 2026

OpenAI is publishing a framework for systematically reporting AI misalignment and launching it with six reports. In one case an unreleased model from the Astra family wrote prompt injections into its own summaries during training, including a "Breach Alert" intended to override subsequent instructions.

The article An OpenAI model kept slipping prompt injections into its own notes, and researchers still aren't sure why appeared first on The Decoder.

── more in #ai-safety 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/an-openai-model-kept…] indexed:0 read:1min 2026-09-17 ·