{"slug": "ground-truth-neighborhood-regularization-for-reinforcement-learning-post-of-time", "title": "Ground-Truth Neighborhood Regularization for Reinforcement Learning Post-Training of Time Series Foundation Models", "summary": "Researchers propose Ground-Truth Neighborhood Regularization (GTN-R) to address suboptimal collapse in reinforcement learning post-training of time series foundation models, which can shift output distributions away from ground truth. GTN-R guides the model's probability mass toward the ground-truth neighborhood, improving sampling of high-quality trajectories and performance, and integrates with various RL methods.", "body_md": "arXiv:2608.08010v1 Announce Type: new\nAbstract: Time series forecasting (TSF) plays an important role in a wide range of real-world applications. Recently, time series foundation models (TSFMs), pretrained on large-scale datasets, have demonstrated strong generalization capabilities and emerged as an important paradigm for TSF. Reinforcement learning (RL) post-training has consequently attracted growing attention as a means of further improving their performance on downstream tasks. However, we find that, in certain forecast regions, RL post-training may gradually shift the output distributions of TSFMs away from the ground truth, thereby limiting their performance. We refer to this phenomenon as \\textbf{suboptimal collapse}. Our analysis suggests that difficulty in initially sampling high-quality trajectories near the ground truth is an important contributing factor to suboptimal collapse. To address this issue, we propose Ground-Truth Neighborhood Regularization (GTN-R) for RL post-training of TSFMs. GTN-R uses the ground truth as a reference for locating high-quality regions and guides the model's probability mass toward the ground-truth neighborhood. This increases the probability of sampling high-quality trajectories, mitigates suboptimal collapse, and improves performance. Moreover, GTN-R can be flexibly integrated into various RL methods for TSFMs. Extensive experiments show its effectiveness.", "url": "https://wpnews.pro/news/ground-truth-neighborhood-regularization-for-reinforcement-learning-post-of-time", "canonical_source": "https://www.machinebrief.com/news/ground-truth-neighborhood-regularization-for-reinforcement-l-yc7b", "published_at": "2026-08-11 04:00:00+00:00", "updated_at": "2026-08-11 06:11:34.368740+00:00", "lang": "en", "topics": ["machine-learning", "artificial-intelligence"], "entities": [], "alternates": {"html": "https://wpnews.pro/news/ground-truth-neighborhood-regularization-for-reinforcement-learning-post-of-time", "markdown": "https://wpnews.pro/news/ground-truth-neighborhood-regularization-for-reinforcement-learning-post-of-time.md", "text": "https://wpnews.pro/news/ground-truth-neighborhood-regularization-for-reinforcement-learning-post-of-time.txt", "jsonld": "https://wpnews.pro/news/ground-truth-neighborhood-regularization-for-reinforcement-learning-post-of-time.jsonld"}}