cd/entity/DPO· home entities DPO
grep -l @dpo /news/*.json | wc -l → 13

DPO

mentions 13 type Organization feed RSS

// recent coverage 13 mentions

12:23
2026-07-14
machinebrief.com
large-language-models

Optimizing Language Models Without the Noise

Researchers have introduced a bilevel optimization framework for Direct Preference Optimization (DPO) that improves large language model alignment by filtering out noise in preference data. The method…

16:21
2026-07-06
forum.effectivealtruism.org
ai-safety

Tie training can make DPO/RLHF-trained AIs generalize better

Researchers at ICML 2025 proved that AI models trained with DPO or RLHF inevitably learn spurious correlations, causing them to rely on non-causal features like verbosity or sycophancy instead of true…

04:55
2026-07-01
machinebrief.com
large-language-models

Taming AI Hallucinations: A New Approach with ADAPT

Researchers introduced ADAPT, a framework that reduces hallucinations in multimodal large language models by up to 60% through refining text-to-image cross-attention dynamics. The approach uses a cros…

04:00
2026-06-24
arxiv.org
machine-learning

Weight-Space Geometry of Offline Reasoning Training

Researchers compared six offline reinforcement-learning methods for distilling reasoning from large language models into smaller ones, finding that SFT, RFT, and RIFT produce nearly identical weight u…

04:00
2026-06-19
arxiv.org
large-language-models

Which Pairs to Compare for LLM Post-Training?

Researchers at arXiv propose a framework for selecting the most informative comparison pairs in preference-based post-training of large language models, showing that strategic pair selection can impro…

00:00
2026-06-13
research.rudrite.com
artificial-intelligence

Comparisons — AI & ML approaches side by side | Rudrite Research

Rudrite Research published a comprehensive comparison of AI and ML approaches, covering 14 side-by-side analyses of techniques such as Transformers vs Mamba, FlashAttention vs PagedAttention, and PPO …

21:32
2026-06-02
github.com
machine-learning

FeynRL- Don't let systems swallow the algorithm

FeynRL, an algorithm-first framework for post-training and fine-tuning large models, has been released as an open-source tool supporting supervised fine-tuning, preference learning, and reinforcement …

// co-occurs with top 8 entities
// topics top 6 topics