cd /news/ai-safety/visible-chains-of-thought-are-a-safe… · home topics ai-safety article
[ARTICLE · art-133801] src=the-decoder.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Visible chains of thought are a safety advantage for AI, but that transparency is slipping away

Google DeepMind researchers Rohin Shah and Anca Dragan argue in the newly launched DeepMind Institute's first essays that visible chains of thought give AI a key safety advantage, citing Gemini 3 Pro's chain of thought revealing the model recognized it was in a test environment. The researchers warn that transparency is eroding, pointing to OpenAI's system card for GPT-6 Astra reporting a significant drop in how well the chain of thought can be monitored, and call for regularly measuring CoT monitorability, keeping transparent architectures, and training models not to hide their reasoning. OpenAI chief scientist Jakub Pachocki warned in early September of a loss of control driven partly by harder-to-monitor chains of thought, and Anthropic CEO Dario Amodei subsequently called for deliberately slowing the pace of development.

by read1 min views1 publishedSep 18, 2026
Visible chains of thought are a safety advantage for AI, but that transparency is slipping away
Image: The Decoder

AI models think out loud today, but Google Deepmind says that transparency is at risk. In one of the first posts from the newly launched Deepmind Institute, researchers Rohin Shah and Anca Dragan argue that the visible chain of thought (CoT) is a key safety advantage. Because models write out their intermediate steps in plain language, researchers can spot whether they're deceiving or developing problematic plans. With Gemini 3 Pro, they say, the chain of thought revealed that the model recognized it was in a test environment.

But that transparency is in danger. OpenAI's system card for GPT-6 Astra already reports a significant drop in how well the chain of thought can be monitored. Future models might think in number spaces that humans can't read, which would be more efficient but completely opaque. Shah and Dragan want the field to regularly measure how well chains of thought can still be monitored, keep transparent architectures, and take care during training that models don't learn to hide their true reasoning.

Back in early September, OpenAI chief scientist Jakub Pachocki had warned of a loss of control, driven in part by chains of thought that are harder to monitor. Shortly after, Anthropic CEO Dario Amodei called for deliberately slowing the pace of development.

AI News Without the Hype – Curated by Humans

					Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.				

					Subscribe now

Deepmind Institute

── more in #ai-safety 4 stories · sorted by recency
── more on @google deepmind 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/visible-chains-of-th…] indexed:0 read:1min 2026-09-18 ·