{"slug": "what-drives-llm-self-reflection-a-controlled-ablation-of-uncertainty-routing-in", "title": "What Drives LLM Self-Reflection? A Controlled Ablation of Uncertainty Routing in Armed Conflict Forecasting", "summary": "A controlled ablation of six conditions in LLM self-reflection for armed conflict forecasting finds that typed action routing, not diagnostic scaffolding or taxonomy vocabulary, drives performance gains, with F1 improving from 0.296 to 0.379 (ΔF1 = +0.075) and a significant overall gain over the single-shot baseline (ΔF1 = +0.101, 95% CI [+0.020, +0.185]). The study, posted on arXiv (2608.12322v1), replicates the vocabulary-routing decomposition on GPT-4o, confirming that action routing provides significant gains (p = 0.025) while taxonomy vocabulary adds no significant value (p = 0.773). Gains concentrate on structurally novel conflicts such as Myanmar (F1: 0.000 → 0.353) and Ukraine (0.167 → 0.500), identifying typed action routing as a promising design principle for metacognitive LLM forecasting agents.", "body_md": "arXiv:2608.12322v1 Announce Type: new\nAbstract: Self-reflection is widely assumed to improve LLM reasoning, yet which component drives the gain remains poorly understood. We present a controlled six-condition ablation isolating four components of LLM self-reflection: evidence exposure, diagnostic scaffolding, taxonomy vocabulary, and action routing. Two precise null results converge on a single mechanism. First, structured diagnostic questions add no measurable value over unstructured reflection ($\\text{F1} = 0.296$ vs $0.297$, $p = 1.000$, 95\\% CI $[-0.041, +0.040]$). Second, presenting the full uncertainty taxonomy while collapsing the action space to a single generic action also adds no value ($\\Delta\\text{F1} = +0.008$, overlapping 95\\% CIs), ruling out taxonomy vocabulary as the mechanism. Typed action routing provides consistent directional gains ($\\text{F1} = 0.379$ vs $0.296$); the conservative estimate controlling for taxonomy vocabulary is $\\Delta\\text{F1} = +0.075$, and the overall gain over the single-shot baseline is significant by bootstrap CI ($\\Delta\\text{F1} = +0.101$, 95\\% CI $[+0.020, +0.185]$). The vocabulary-routing decomposition replicates on GPT-4o: taxonomy vocabulary adds no significant value over generic reflection ($p = 0.773$), while action routing provides significant gains ($p = 0.025$), confirming the mechanism holds across backbones. Gains concentrate on structurally novel conflicts: in Myanmar ($\\text{F1}: 0.000 \\rightarrow 0.353$) and Ukraine ($0.167 \\rightarrow 0.500$), the vocabulary-only condition recovers no more than generic reflection while action routing breaks the degenerate prior. These findings identify typed action routing -- not diagnostic scaffolding or taxonomy vocabulary -- as a promising design principle for metacognitive LLM forecasting agents, while motivating larger-scale evaluation across conflict typologies.", "url": "https://wpnews.pro/news/what-drives-llm-self-reflection-a-controlled-ablation-of-uncertainty-routing-in", "canonical_source": "https://arxiv.org/abs/2608.12322", "published_at": "2026-08-14 04:00:00+00:00", "updated_at": "2026-08-14 04:12:07.221759+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-agents"], "entities": ["arXiv", "GPT-4o", "Myanmar", "Ukraine"], "alternates": {"html": "https://wpnews.pro/news/what-drives-llm-self-reflection-a-controlled-ablation-of-uncertainty-routing-in", "markdown": "https://wpnews.pro/news/what-drives-llm-self-reflection-a-controlled-ablation-of-uncertainty-routing-in.md", "text": "https://wpnews.pro/news/what-drives-llm-self-reflection-a-controlled-ablation-of-uncertainty-routing-in.txt", "jsonld": "https://wpnews.pro/news/what-drives-llm-self-reflection-a-controlled-ablation-of-uncertainty-routing-in.jsonld"}}