cd /news/large-language-models/metadata-free-meta-reweighted-direct… · home topics large-language-models article
[ARTICLE · art-58277] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels

A new study proposes Metadata-Free Meta-Reweighted Direct Preference Optimization (DPO) to improve alignment of large language models under noisy preference labels. The method uses a bilevel optimization framework and a task-agnostic meta-knowledge-driven approach to handle noisy data without requiring metadata. Experiments on TL;DR summarization and Anthropic HH single-turn dialogue show improved performance over DPO baselines under different noise rates.

read1 min views13 publishedJul 14, 2026

arXiv:2607.09796v1 Announce Type: new Abstract: Direct Preference Optimization (DPO) has become an important method for aligning large language models (LLMs) with human preferences because it removes the need for explicit reward modeling and reinforcement learning optimization. However, its performance depends heavily on the quality of preference data, and noisy preference data in real-world settings can weaken alignment performance. To address this issue, we propose a bilevel optimization framework and prove, under certain assumptions, that this framework can recover the DPO optimum under clean data. We further derive a prior form for the learnable weighting function under asymmetric label-flipping noise. Considering that high-quality metadata may be difficult to obtain, we propose a task-agnostic meta-knowledge-driven method that enables meta-learning even when metadata is completely unavailable. To reduce the high cost of higher-order gradients in LLM meta-learning, we combine central-difference approximation with LoRA fine-tuning and develop a scalable training scheme. Experiments on TL;DR summarization and Anthropic HH single-turn dialogue show that the proposed method improves training performance over multiple DPO baselines under different noise rates.

── more in #large-language-models 4 stories · sorted by recency
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/metadata-free-meta-r…] indexed:0 read:1min 2026-07-14 ·