{"slug": "s2t-rlhf-hierarchical-credit-assignment-for-stable-preference-based-rlhf", "title": "S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF", "summary": "Researchers propose S2T-RLHF, a sentence-to-token reward decomposition framework that improves training stability in reinforcement learning from human feedback (RLHF) by allocating sequence-level preference rewards across sentences before applying bounded token-level refinement. The method, detailed in arXiv:2607.18258v1, addresses instability caused by noisy preference signals and ambiguous credit assignment, outperforming standard RLHF across multiple datasets without requiring reward-model retraining or token-level supervision.", "body_md": "arXiv:2607.18258v1 Announce Type: new\nAbstract: Reinforcement learning from human feedback (RLHF) with preference-based reward models often exhibits unstable training dynamics. A key contributing factor is that standard RLHF relies on a single sequence-level scalar reward, which is propagated to token-level policy updates and leaves credit assignment within a response inherently ambiguous. Recent work has attempted to address this issue by refining rewards into denser token-level supervision, often relying on the implicit assumption that finer-grained credit assignment improves optimization. We argue that this assumption is incomplete: when preference signals are noisy and only defined at the response level, overly fine-grained reward refinement can amplify reward uncertainty and destabilize learning. To address this problem, we propose a granularity-aware principle for hierarchical credit assignment, emphasizing stability-oriented reward design rather than maximal allocation precision. Under this principle, sentences serve as a natural intermediate granularity, balancing semantic coherence with robustness to token-level noise. Guided by this view, we introduce S2T-RLHF. This sentence-to-token reward decomposition framework first allocates sequence-level preference rewards across sentences and then applies bounded token-level refinement within each sentence, without reward-model retraining or token-level supervision. Experiments across multiple datasets and optimization settings show that S2T-RLHF improves training stability and robustness while maintaining competitive preference alignment.", "url": "https://wpnews.pro/news/s2t-rlhf-hierarchical-credit-assignment-for-stable-preference-based-rlhf", "canonical_source": "https://arxiv.org/abs/2607.18258", "published_at": "2026-07-22 04:00:00+00:00", "updated_at": "2026-07-22 04:12:29.232477+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-research"], "entities": ["arXiv"], "alternates": {"html": "https://wpnews.pro/news/s2t-rlhf-hierarchical-credit-assignment-for-stable-preference-based-rlhf", "markdown": "https://wpnews.pro/news/s2t-rlhf-hierarchical-credit-assignment-for-stable-preference-based-rlhf.md", "text": "https://wpnews.pro/news/s2t-rlhf-hierarchical-credit-assignment-for-stable-preference-based-rlhf.txt", "jsonld": "https://wpnews.pro/news/s2t-rlhf-hierarchical-credit-assignment-for-stable-preference-based-rlhf.jsonld"}}