cd /news/ai-safety/first-token-matters-understanding-sa… · home topics ai-safety article
[ARTICLE · art-132295] src=machinebrief.com ↗ pub= topic=ai-safety verified=true sentiment=· neutral

First Token Matters: Understanding Safety Collapse in Large Reasoning Models

A new arXiv paper (2609.18471v1) identifies a failure mode it calls Onset Refusal Collapse (ORC), in which the refusal-related signal of Large Reasoning Models drops sharply at the first generated token under harmful queries, correlating with unsafe response generation. The authors propose SafeToken, a lightweight inference-time intervention that injects a learned continuous safety anchor at reasoning onset; updating only a single token embedding, SafeToken mitigates ORC, improves safety on harmful-query benchmarks, and largely preserves reasoning utility. The work attributes safety failures in LRMs to a transient breakdown at the transition from understanding to generation rather than to a lack of training.

by read1 min views1 publishedSep 17, 2026

arXiv:2609.18471v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) exhibit strong problem-solving abilities, yet their safety alignment often degrades when handling harmful queries. Existing approaches to improving safety largely rely on additional training or preference optimization, while offering limited understanding of the internal mechanisms behind safety failures. In this work, we investigate this failure through a token-level positional analysis of refusal dynamics and identify a localized vulnerability at the onset of reasoning, which we term Onset Refusal Collapse (ORC). We find that the refusal-related signal of LRMs drops sharply at the first generated token under harmful queries, which is associated with unsafe response generation. Motivated by this finding, we propose SafeToken, a lightweight inference-time intervention that injects a learned continuous safety anchor precisely at reasoning onset. Despite updating only a single token embedding, SafeToken effectively mitigates ORC, improves safety on harmful-query benchmarks, and largely preserves reasoning utility. These results suggest that safety failures in LRMs can arise from a transient breakdown at the critical transition from understanding to generation.

── more in #ai-safety 4 stories · sorted by recency
── more on @large reasoning models 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/first-token-matters-…] indexed:0 read:1min 2026-09-17 ·