{"slug": "gradient-aligned-pair-selection-for-personalized-preference-optimization", "title": "Gradient-Aligned Pair Selection for Personalized Preference Optimization", "summary": "Researchers posted arXiv:2610.00061v1, introducing GAP-DPO (Geometry-Aligned Preference DPO), an iterative algorithm that selects preference pairs by aligning their gradient direction with expected user utility gradients while controlling distribution shift through epoch-wise regeneration. The authors formalize personalized preference learning as a geometry-aligned optimization problem, showing that under off-policy sampling the DPO update shifts from a purely error-corrective signal to a reinforcement-like update when preference margins align directionally with utility gradients. Experiments on personalized text generation benchmarks show GAP-DPO consistently improves stylistic fidelity, preference alignment, and generation quality over standard DPO variants.", "body_md": "arXiv:2610.00061v1 Announce Type: new \nAbstract: Personalizing large language models (LLMs) requires aligning generation behavior with user-specific preferences rather than aggregate quality. While Direct Preference Optimization (DPO) provides a stable framework for preference learning, its effectiveness in personalized settings critically depends on how preference pairs are selected. Existing approaches typically rely on heuristic criteria, such as likelihood-based extremes, which decouple optimization from explicit user utility and can lead to degraded personalization. We formalize personalized preference learning as a geometry-aligned optimization problem by analyzing the first-order interaction between gradients of expected user utility and DPO update directions. Our analysis reveals that, under off-policy sampling, the DPO update transitions from a purely error-corrective signal to a reinforcement-like update when preference margins are directionally aligned with utility gradients. This perspective exposes pair selection as a geometric decision that governs whether preference optimization advances or hinders personalization. Motivated by this insight, we propose GAP-DPO (Geometry-Aligned Preference DPO), an iterative algorithm that performs utility-aware, geometry-aligned pair selection while controlling distribution shift via epoch-wise regeneration. Experiments on personalized text generation benchmarks show that GAP-DPO consistently improves stylistic fidelity, preference alignment, and generation quality compared to standard DPO variants. Together, our results establish gradient alignment as a unifying principle for personalized preference optimization and demonstrate that pair selection is an intrinsic component of the optimization geometry rather than a heuristic preprocessing step.", "url": "https://wpnews.pro/news/gradient-aligned-pair-selection-for-personalized-preference-optimization", "canonical_source": "https://arxiv.org/abs/2610.00061", "published_at": "2026-10-02 04:00:00+00:00", "updated_at": "2026-10-02 04:15:13.546053+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research"], "entities": ["GAP-DPO", "Direct Preference Optimization", "arXiv:2610.00061v1"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/gradient-aligned-pair-selection-for-personalized-preference-optimization", "markdown": "https://wpnews.pro/news/gradient-aligned-pair-selection-for-personalized-preference-optimization.md", "text": "https://wpnews.pro/news/gradient-aligned-pair-selection-for-personalized-preference-optimization.txt", "jsonld": "https://wpnews.pro/news/gradient-aligned-pair-selection-for-personalized-preference-optimization.jsonld"}}