{"slug": "length-value-model-scalable-value-pretraining-for-token-level-length-modeling", "title": "Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling", "summary": "Researchers from University of California, Santa Barbara, Carnegie Mellon University, LMSYS Org, and University of Wisconsin–Madison introduced the Length Value Model (LenVM), a token-level framework that models remaining generation length at each decoding step. On the LIFEBench exact length matching task, applying LenVM to a 7B model improved the length score from 30.9 to 64.8, outperforming frontier closed-source models. LenVM enables continuous control over the trade-off between performance and efficiency, maintaining 63% accuracy on GSM8K at a 200-token budget compared to 6% for a token budget baseline.", "body_md": "Token serves as the fundamental unit of computation in modern autoregressive models, and generation length directly influences both inference cost and reasoning performance. Despite its importance, existing approaches lack fine-grained length modeling, operating primarily at the coarse-grained sequence level. In this paper, we introduce the Length Value Model (LenVM), a token-level framework that models the remaining generation length at each decoding step. By formulating length modeling as a value estimation problem and assigning a constant negative reward to each generated token, LenVM predicts a bounded, discounted return that serves as a monotone proxy for the remaining generation horizon. This formulation yields supervision that is annotation-free, dense, unbiased, and scalable. Experiments on LLMs and VLMs demonstrate that LenVM provides a highly effective signal at inference time. On the LIFEBench exact length matching task, applying LenVM to a 7B model improves the length score from 30.9 to 64.8, significantly outperforming frontier closed-source models. Furthermore, LenVM enables continuous control over the trade off between performance and efficiency. On GSM8K at a budget of 200 tokens, LenVM maintains 63 percent accuracy compared to 6 percent for token budget baseline. It also accurately predicts total generation length from the prompt boundary. Finally, LenVM’s token-level values offer an interpretable view of generation dynamics, revealing how specific tokens shift reasoning toward shorter or longer regimes. Results demonstrate that LenVM supports a broad range of applications, including length control, prediction, and interpretation of generation dynamics. They suggest that generation length can be effectively modeled as a token-level value signal, highlighting the potential of LenVM as a general framework for length modeling and as a length-specific value signal that could support future RL training.\n\n- † University of California, Santa Barbara\n- ‡ Carnegie Mellon University\n- § LMSYS Org\n- ¶ University of Wisconsin–Madison", "url": "https://wpnews.pro/news/length-value-model-scalable-value-pretraining-for-token-level-length-modeling", "canonical_source": "https://machinelearning.apple.com/research/length-value-model", "published_at": "2026-07-20 00:00:00+00:00", "updated_at": "2026-07-20 22:40:06.577217+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research"], "entities": ["University of California, Santa Barbara", "Carnegie Mellon University", "LMSYS Org", "University of Wisconsin–Madison", "Length Value Model", "LenVM", "LIFEBench", "GSM8K"], "alternates": {"html": "https://wpnews.pro/news/length-value-model-scalable-value-pretraining-for-token-level-length-modeling", "markdown": "https://wpnews.pro/news/length-value-model-scalable-value-pretraining-for-token-level-length-modeling.md", "text": "https://wpnews.pro/news/length-value-model-scalable-value-pretraining-for-token-level-length-modeling.txt", "jsonld": "https://wpnews.pro/news/length-value-model-scalable-value-pretraining-for-token-level-length-modeling.jsonld"}}