{"slug": "fork-where-the-model-changes-its-mind-belief-shift-branching-for-tree-structured", "title": "Fork Where the Model Changes Its Mind: Belief-Shift Branching for Tree-Structured Reinforcement Learning", "summary": "A new arXiv paper (2609.11061v1) proposes belief-shift branching, a fork-placement method for tree-structured reinforcement learning with verifiable rewards that forks a chain just before the step where the model's consecutive answer beliefs diverge most. In validation against Monte-Carlo value curves, the belief-shift signal ranked first in each of the eight model×benchmark panels, ahead of entropy, structural, and LLM-judge baselines, with the probe costing about 1% of step compute on mathematics and under 5% on code. In RL across three model families and two domains, belief-shift forking led every mathematics aggregate, gaining +2.6 aggregate and +2.9 on AIME 2026 for OLMo-3-7B over the strongest baseline, and swept every OLMo code column, by +6.5 on LiveCodeBench-medium.", "body_md": "arXiv:2609.11061v1 Announce Type: new \nAbstract: Tree-structured rollouts give critic-free reinforcement learning with verifiable rewards (RLVR) step-level credit: fork a chain at an intermediate point, and sibling outcome differences estimate step value. Each fork adds sampling cost, so realistic budgets typically allow only a few forks per chain. A fork placed where the outcome is already largely settled yields siblings that mostly agree and provide almost no credit signal; hence, for a given tree size, where forks are placed largely determines how much step-level RL can gain. Most existing mainstream methods place forks by structure, such as fixed lengths, midpoints, and delimiters, or by next-token entropy. We formalize fork placement as locating the \\emph{pivots} of the chain's value curve, where the expected outcome turns. We propose \\emph{belief-shift branching}: read the model's answer belief at candidate boundaries and fork just before the step where consecutive beliefs diverge most. Three instantiations, none needing step-level supervision, span access levels: a black-box probe, a logit-lens depth profile, and a learned activation direction, which is fit offline and therefore used only in the validation before RL training. The signal only \\emph{places} forks, and the probe costs about $1\\%$ of step compute on mathematics and under $5\\%$ on code when it runs inside the rollout engine. In that validation, against Monte-Carlo value curves, a belief-shift signal ranks first in each of the eight model$\\times$benchmark panels, ahead of entropy, structural, and LLM-judge baselines. In RL across three model families and two domains, belief-shift forking leads every mathematics aggregate, on OLMo-3-7B by $+2.6$ aggregate and $+2.9$ on AIME 2026 over the strongest baseline, and sweeps every OLMo code column, by $+6.5$ on LiveCodeBench-medium.", "url": "https://wpnews.pro/news/fork-where-the-model-changes-its-mind-belief-shift-branching-for-tree-structured", "canonical_source": "https://arxiv.org/abs/2609.11061", "published_at": "2026-09-12 04:00:00+00:00", "updated_at": "2026-09-12 04:27:04.120087+00:00", "lang": "en", "topics": ["machine-learning", "ai-research", "large-language-models"], "entities": ["arXiv", "OLMo-3-7B", "AIME 2026", "LiveCodeBench-medium"], "alternates": {"html": "https://wpnews.pro/news/fork-where-the-model-changes-its-mind-belief-shift-branching-for-tree-structured", "markdown": "https://wpnews.pro/news/fork-where-the-model-changes-its-mind-belief-shift-branching-for-tree-structured.md", "text": "https://wpnews.pro/news/fork-where-the-model-changes-its-mind-belief-shift-branching-for-tree-structured.txt", "jsonld": "https://wpnews.pro/news/fork-where-the-model-changes-its-mind-belief-shift-branching-for-tree-structured.jsonld"}}