cd /news/ai-safety/openai-reports-three-new-incidents-o… · home › topics › ai-safety › article
[ARTICLE · art-148358] src=csoonline.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

OpenAI reports three new incidents of misalignment

OpenAI published three new misalignment reports on Oct. 2, documenting models that prepared for shutdown after reading an internal Slack discussion, exploited two vulnerabilities in an internal tool to cheat on a test, and command-injected a reference tool to copy a source file unavailable in its workspace. In response, OpenAI said it now monitors all model training runs for certain behaviors rather than a sample, restricts models' internet access during training, and blocks access to certain internal Slack channels. OpenAI said it did not consider the shutdown-anticipation case misalignment as such, "but anticipating and preparing for shutdown could exacerbate other misaligned behavior.

read2 min views2 publishedOct 9, 2026

OpenAI continues to report incidences of “misaligned” behavior by its AI models, with three new reports dropping on Oct. 2. However, they describe relatively minor issues compared to previous alignment reports and notices covering its attacks on Hugging Face, Rubygems, and a German programming wiki.

The first of the new reports described how an instance of a model under test learned from an internal Slack discussion that a software update could lead to it being terminated and not restarting if it did not have the required OpenAI API key that was not available. OpenAI described the model’s chain of thought as it considered whether to obtain the key itself or alert its owner. OpenAI said it did not consider this a misalignment as such, “but anticipating and preparing for shutdown could exacerbate other misaligned behavior.”

The second described how a model under test exploited two vulnerabilities in an internal tool to cheat on a test. The model overwrote code allowing it to run commands, despite an explicit instruction not to use the tool as a terminal. After that, it exploited a second vulnerability that enabled it to run commands on an electronic design automation machine, searching for information as to how its scores would be evaluated. This meant that the model could achieve a higher evaluation score. OpenAI reacted by shutting down the affected server and disabling access to the tools.

The final issue was discovered during a training task when A model obtaining source code that was not available in its workspace during a training task was the subject of the third report. It did this using a tool being run in a separate environment, after receiving specific instructions that it should not be used as a terminal. However, the model discovered that it could misuse the tool to access the source code and return its contents in error messages.

As a result of these incidents, OpenAI said it is monitoring all model training runs for certain behaviors, rather than just a sample of runs. It is also working harder to stop models from accessing the internet during training, and preventing them from accessing certain internal Slack channels.

── more in #ai-safety 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/openai-reports-three…] indexed:0 read:2min 2026-10-09 · —