{"slug": "data-dpo-direct-preference-optimization-for-target-model-data-selection-in-llm", "title": "Data-DPO: Direct Preference Optimization for Target Model Data Selection in LLM Post-Training", "summary": "Researchers propose Data-DPO, a target model-oriented data selection method for supervised fine-tuning that uses one-step probing to capture target-model-aware data preferences and combines them with external quality scores and marginal diversity. Experiments on Vision-Flan and LLaVA-CoT show Data-DPO consistently outperforms existing baselines across multiple data budgets and surpasses full data training performance.", "body_md": "arXiv:2608.16926v1 Announce Type: new\nAbstract: Data selection in supervised fine-tuning aims to select a small set of effective samples from large-scale candidate data, reducing training cost while preserving model performance. However, existing methods usually treat data value as a relatively static property, and pay limited attention to the compatibility between data and the capability distribution of the target model. To address this issue, we propose Data-DPO, a target model-oriented SFT data selection method. Data-DPO observes the local training feedback of the target model on different samples through one-step probing, transforms activation differences among samples into pairwise data preferences, and trains a lightweight reward model to learn target-model-aware data preferences. In the final selection stage, Data-DPO further combines target model preference, external quality scores, and marginal diversity to construct a more stable and effective training subset. Experimental results on Vision-Flan and LLaVA-CoT show that Data-DPO consistently outperforms existing data selection baselines under multiple data budgets and stably surpasses full data training performance.", "url": "https://wpnews.pro/news/data-dpo-direct-preference-optimization-for-target-model-data-selection-in-llm", "canonical_source": "https://arxiv.org/abs/2608.16926", "published_at": "2026-08-19 04:00:00+00:00", "updated_at": "2026-08-19 04:13:17.995159+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research"], "entities": ["Data-DPO", "Vision-Flan", "LLaVA-CoT"], "alternates": {"html": "https://wpnews.pro/news/data-dpo-direct-preference-optimization-for-target-model-data-selection-in-llm", "markdown": "https://wpnews.pro/news/data-dpo-direct-preference-optimization-for-target-model-data-selection-in-llm.md", "text": "https://wpnews.pro/news/data-dpo-direct-preference-optimization-for-target-model-data-selection-in-llm.txt", "jsonld": "https://wpnews.pro/news/data-dpo-direct-preference-optimization-for-target-model-data-selection-in-llm.jsonld"}}