{"slug": "the-knowing-saying-gap-when-probes-see-errors-that-confidence-misses", "title": "The Knowing-Saying Gap: When Probes See Errors that Confidence Misses", "summary": "A new study from arXiv (2608.07528v1) finds that linear probes detect corrupted context in language models with near-perfect accuracy but fail to predict final answer correctness, revealing a 'knowing-saying gap' across multi-hop arithmetic chains and model families. The research, which refutes the pre-registered 'persistence beats peak' hypothesis, shows probe-based interventions are model- and error-type-dependent, with branch-and-pick being net-positive across models and uniquely non-breaking on Llama-3.1-8B (4 rescued, 0 broken), while reprompt and replace-prior break correct traces at roughly the rate they rescue wrong ones.", "body_md": "arXiv:2608.07528v1 Announce Type: new\nAbstract: Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure prediction. The result is a dissociation with direct implications for deployment monitoring. Across multi-hop arithmetic chains, probes that detect corruption turn out to be uninformative about final answer correctness; models forced into structured confidence formats collapse to two values with indistinguishable error rates; and probe persistence across hops fails to separate correct from incorrect outcomes, refuting our pre-registered \"persistence beats peak\" hypothesis. This pattern of knowing but not saying generalises across model families including reasoning models. As a real-time monitor, probe-based interventions are sharply model and error-type dependent: branch-and-pick is net-positive across models and uniquely non-breaking on Llama-3.1-8B (4 rescued, 0 broken), while reprompt and replace-prior break correct traces at roughly the rate they rescue wrong ones. Probe-based monitoring is a necessary complement to verbalised confidence, but no single intervention dominates, and the deployable answer is model-aware, error-type-aware routing.", "url": "https://wpnews.pro/news/the-knowing-saying-gap-when-probes-see-errors-that-confidence-misses", "canonical_source": "https://arxiv.org/abs/2608.07528", "published_at": "2026-08-11 04:00:00+00:00", "updated_at": "2026-08-11 04:16:08.577308+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-research", "ai-safety"], "entities": ["arXiv", "Llama-3.1-8B"], "alternates": {"html": "https://wpnews.pro/news/the-knowing-saying-gap-when-probes-see-errors-that-confidence-misses", "markdown": "https://wpnews.pro/news/the-knowing-saying-gap-when-probes-see-errors-that-confidence-misses.md", "text": "https://wpnews.pro/news/the-knowing-saying-gap-when-probes-see-errors-that-confidence-misses.txt", "jsonld": "https://wpnews.pro/news/the-knowing-saying-gap-when-probes-see-errors-that-confidence-misses.jsonld"}}