{"slug": "meta-learned-reward-shaping-for-reinforcement-learning-from-human-feedback", "title": "Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback", "summary": "Researchers introduce MeRLa (Meta-Learned Reward Shaping), a framework that meta-learns a task-aware shaping function for Reinforcement Learning from Human Feedback (RLHF) to improve alignment of large language models. Experiments on LLaMA-3-8B show MeRLa achieves a 90.8% length-controlled win rate on AlpacaEval 2.0 and a score of 9.14 on MT-Bench, with 41% less training instability compared to PPO, DPO, GRPO, and DAPO.", "body_md": "arXiv:2607.26094v1 Announce Type: new\nAbstract: Reinforcement Learning from Human Feedback (RLHF) is the standard approach for aligning large language models with human preferences, but its quality is limited by static, task-agnostic reward models. This mismatch leads to sparse learning signals and suboptimal alignment. We introduce MeRLa (Meta-Learned Reward Shaping), a principled framework that meta-learns a task-aware shaping function $\\Phi(x,y;\\phi)$ across auxiliary tasks before RLHF training. The learned shaping produces a composite reward that preserves policy optimality while providing task-specific learning signals. Our meta-objective combines task discrimination, entropy regularization, and potential-based conservation for stable convergence. We provide theoretical guarantees for policy invariance, analyze representation drift sensitivity, and formally address incentive misalignment from entropy maximization. Experiments on LLaMA-3-8B across four benchmarks show consistent improvements over PPO, DPO, GRPO, and DAPO, achieving a 90.8% length-controlled win rate on AlpacaEval 2.0 and a score of 9.14 on MT-Bench, with 41% less training instability. MeRLa retains its benefits when combined with process-based and rubric-based enhanced rewards.", "url": "https://wpnews.pro/news/meta-learned-reward-shaping-for-reinforcement-learning-from-human-feedback", "canonical_source": "https://arxiv.org/abs/2607.26094", "published_at": "2026-07-30 04:00:00+00:00", "updated_at": "2026-07-30 04:30:01.508783+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research"], "entities": ["MeRLa", "LLaMA-3-8B", "AlpacaEval 2.0", "MT-Bench", "PPO", "DPO", "GRPO", "DAPO"], "alternates": {"html": "https://wpnews.pro/news/meta-learned-reward-shaping-for-reinforcement-learning-from-human-feedback", "markdown": "https://wpnews.pro/news/meta-learned-reward-shaping-for-reinforcement-learning-from-human-feedback.md", "text": "https://wpnews.pro/news/meta-learned-reward-shaping-for-reinforcement-learning-from-human-feedback.txt", "jsonld": "https://wpnews.pro/news/meta-learned-reward-shaping-for-reinforcement-learning-from-human-feedback.jsonld"}}