04:00
2026-08-19
arxiv.org
artificial-intelligence
Data-DPO: Direct Preference Optimization for Target Model Data Selection in LLM Post-Training
Researchers propose Data-DPO, a target model-oriented data selection method for supervised fine-tuning that uses one-step probing to capture target-model-aware data preferences and combines them with โฆ