cd /news/large-language-models/targeted-retrieval-compact-represent… · home › topics › large-language-models › article
[ARTICLE · art-143033] src=machinebrief.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Targeted Retrieval, Compact Representations: How CoT Reasoning Improves Long-Context Counting

A new arXiv paper (2609.38958v1) reports that Chain-of-Thought reasoning improves long-context counting accuracy across twelve model comparison groups, with the largest gains at higher counts. The authors' mechanistic analysis of a needle-in-a-haystack counting task identifies two contrasting mechanisms: Non-thinking models use broad retrieval, attending to multiple needles at once, while Thinking models use targeted retrieval, enumerating needles successively in CoT traces and concentrating attention on individual needles with more compact internal representations. Causal intervention analysis suggests Thinking models use the CoT trace to maintain and update an internal counter as needles are retrieved, even without explicit numbering.

by read1 min views1 publishedOct 1, 2026

arXiv:2609.38958v1 Announce Type: new Abstract: Large language models (LLMs) have been rapidly improving in long-context tasks, powered by Chain-of-Thought (CoT) reasoning. However, the internal mechanisms underlying this improvement remain unclear. We investigate these mechanisms through a needle-in-a-haystack (NIAH) counting task, where an LLM is asked to count the number of records dispersed in a long text. Across twelve model comparison groups, Thinking (or reasoning) improves counting accuracy over Non-thinking, with pronounced gains at larger counts. This motivates our mechanistic analysis, which identifies two contrasting mechanisms: (i) broad retrieval, where Non-thinking models broadly attend to multiple needles; (ii) targeted retrieval, where Thinking models use enumeration in CoT traces to successively retrieve needles. Targeted retrieval concentrates attention on individual needles and is accompanied by more compact internal representations. Moreover, causal intervention analysis suggests that Thinking models use the CoT trace to maintain and update an internal counter as needles are successively retrieved, even without explicit numbering. In small controlled experiments, both retrieval mechanisms and counter states emerge under standard autoregressive training. Together, our results connect long-context retrieval with representation geometry of counting, supporting a state-tracking account of CoT reasoning.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/targeted-retrieval-c…] indexed:0 read:1min 2026-10-01 · —