{"slug": "enabling-preference-driven-unlearning-in-few-step-distilled-text-to-image-models", "title": "Enabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion Models", "summary": "A new arXiv paper (2610.10859v1) introduces a preference-driven unlearning framework that removes targeted concepts such as identity and NSFW content directly from few-step distilled (FSD) text-to-image diffusion models. The authors report that standard Direct Preference Optimization (DPO) and its unlearning derivatives, which are formulated around noise-prediction error, transfer poorly to FSD models because of their altered generation dynamics, and that re-distilling an unlearned base model instead incurs substantial computational and time overhead. Their modified preference optimization formulation is explicitly aligned with few-step generation properties, and experiments show consistent forgetting with strong retention of non-targeted capabilities, with the method also extended to object-level unlearning.", "body_md": "arXiv:2610.10859v1 Announce Type: new \nAbstract: Text-to-image diffusion models are increasingly distilled into few-step variants and being deployed to enable fast inference. However, their ability to generate harmful or undesired content poses significant safety risks. Data-driven unlearning methods suppress targeted generations by fine-tuning model weights using specialized unlearning objectives. Crucially, these objectives implicitly rely on multi-step denoising dynamics, an assumption that breaks down for few-step distilled (FSD) models, resulting in ineffective forgetting. Furthermore, performing unlearning on the non-distilled base model and subsequently re-distilling it to obtain an unlearned FSD model incurs substantial computational and time overhead, making it impractical in many settings. Hence, we address this limitation with a preference-driven unlearning framework that revisits Direct Preference Optimization (DPO) for diffusion models. We show that standard DPO and its unlearning derivatives, formulated around noise-prediction error, transfer poorly to FSD models due to their altered generation dynamics. To overcome this, we introduce a modified preference optimization formulation explicitly aligned with the few-step generation properties, enabling direct concept removal in FSD models while preserving few-step efficiency and maintaining strong retention of desirable (non-targeted) capabilities. We evaluate our framework primarily on identity and NSFW (nudity) removal tasks and also extend our method to object-level unlearning. Extensive experiments demonstrate consistent and effective forgetting, and strong retention performance, establishing our method as a practical and principled solution for unlearning in FSD models.", "url": "https://wpnews.pro/news/enabling-preference-driven-unlearning-in-few-step-distilled-text-to-image-models", "canonical_source": "https://arxiv.org/abs/2610.10859", "published_at": "2026-10-09 04:00:00+00:00", "updated_at": "2026-10-09 04:17:05.107235+00:00", "lang": "en", "topics": ["artificial-intelligence", "generative-ai", "ai-safety", "ai-research", "machine-learning"], "entities": ["arXiv", "Direct Preference Optimization", "few-step distilled diffusion models"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/enabling-preference-driven-unlearning-in-few-step-distilled-text-to-image-models", "markdown": "https://wpnews.pro/news/enabling-preference-driven-unlearning-in-few-step-distilled-text-to-image-models.md", "text": "https://wpnews.pro/news/enabling-preference-driven-unlearning-in-few-step-distilled-text-to-image-models.txt", "jsonld": "https://wpnews.pro/news/enabling-preference-driven-unlearning-in-few-step-distilled-text-to-image-models.jsonld"}}