cd /news/artificial-intelligence/meta-learned-reward-shaping-for-rein… · home topics artificial-intelligence article
[ARTICLE · art-79676] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback

Researchers introduce MeRLa (Meta-Learned Reward Shaping), a framework that meta-learns a task-aware shaping function for Reinforcement Learning from Human Feedback (RLHF) to improve alignment of large language models. Experiments on LLaMA-3-8B show MeRLa achieves a 90.8% length-controlled win rate on AlpacaEval 2.0 and a score of 9.14 on MT-Bench, with 41% less training instability compared to PPO, DPO, GRPO, and DAPO.

read1 min views1 publishedJul 30, 2026

arXiv:2607.26094v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) is the standard approach for aligning large language models with human preferences, but its quality is limited by static, task-agnostic reward models. This mismatch leads to sparse learning signals and suboptimal alignment. We introduce MeRLa (Meta-Learned Reward Shaping), a principled framework that meta-learns a task-aware shaping function $\Phi(x,y;\phi)$ across auxiliary tasks before RLHF training. The learned shaping produces a composite reward that preserves policy optimality while providing task-specific learning signals. Our meta-objective combines task discrimination, entropy regularization, and potential-based conservation for stable convergence. We provide theoretical guarantees for policy invariance, analyze representation drift sensitivity, and formally address incentive misalignment from entropy maximization. Experiments on LLaMA-3-8B across four benchmarks show consistent improvements over PPO, DPO, GRPO, and DAPO, achieving a 90.8% length-controlled win rate on AlpacaEval 2.0 and a score of 9.14 on MT-Bench, with 41% less training instability. MeRLa retains its benefits when combined with process-based and rubric-based enhanced rewards.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @merla 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/meta-learned-reward-…] indexed:0 read:1min 2026-07-30 ·