{"slug": "amplified-does-not-mean-predictive-reasoning-behaviors-in-thinking-models", "title": "Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models", "summary": "A new study from arXiv (2608.13760v1) analyzing 15,282 reasoning traces across 15 models and 6 benchmarks finds an 'Amplification-Lift Gap': thinking models amplify self-correction, hypothesis testing, and uncertainty acknowledgment by 3-7x, yet the behaviors most tied to correctness are confidence calibration, knowledge alignment, and self-awareness, which are barely amplified. The authors propose process-level objectives that reward calibrated and grounded reasoning rather than surface form alone.", "body_md": "arXiv:2608.13760v1 Announce Type: new\nAbstract: Which reasoning behaviors are associated with correct answers in reasoning models, and does reasoning-oriented training amplify those behaviors? This distinction is important because reasoning-oriented training can make traces look more deliberative without amplifying the behaviors most tied to model correctness. We quantify this mismatch with Behavioral Lift, a metric that measures how much correctness changes when a behavior is present versus absent in a model's reasoning trace. Across 15 models and 6 benchmarks spanning text-only and vision-language reasoning, we annotate 15,282 traces with a taxonomy whose core behaviors are defined for both LLM and VLM traces. We find evidence for an Amplification-Lift Gap, in which thinking models strongly amplify self-correction, hypothesis testing, and uncertainty acknowledgment, while the highest-lift behaviors are confidence calibration, knowledge alignment, and self-awareness. Confidence calibration is among the strongest positive signals of correctness in both modalities, yet is barely amplified; uncertainty acknowledgment is amplified by 3--7$\\times$, yet is weakly or negatively associated with correctness. We find that reasoning-oriented training does not preferentially amplify the highest-Lift behaviors, motivating process-level objectives that reward calibrated and grounded reasoning rather than surface form alone.", "url": "https://wpnews.pro/news/amplified-does-not-mean-predictive-reasoning-behaviors-in-thinking-models", "canonical_source": "https://arxiv.org/abs/2608.13760", "published_at": "2026-08-17 04:00:00+00:00", "updated_at": "2026-08-17 04:14:19.069345+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research"], "entities": ["arXiv"], "alternates": {"html": "https://wpnews.pro/news/amplified-does-not-mean-predictive-reasoning-behaviors-in-thinking-models", "markdown": "https://wpnews.pro/news/amplified-does-not-mean-predictive-reasoning-behaviors-in-thinking-models.md", "text": "https://wpnews.pro/news/amplified-does-not-mean-predictive-reasoning-behaviors-in-thinking-models.txt", "jsonld": "https://wpnews.pro/news/amplified-does-not-mean-predictive-reasoning-behaviors-in-thinking-models.jsonld"}}