cd /news/artificial-intelligence/enabling-preference-driven-unlearnin… · home › topics › artificial-intelligence › article
[ARTICLE · art-148009] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Enabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion Models

A new arXiv paper (2610.10859v1) introduces a preference-driven unlearning framework that removes targeted concepts such as identity and NSFW content directly from few-step distilled (FSD) text-to-image diffusion models. The authors report that standard Direct Preference Optimization (DPO) and its unlearning derivatives, which are formulated around noise-prediction error, transfer poorly to FSD models because of their altered generation dynamics, and that re-distilling an unlearned base model instead incurs substantial computational and time overhead. Their modified preference optimization formulation is explicitly aligned with few-step generation properties, and experiments show consistent forgetting with strong retention of non-targeted capabilities, with the method also extended to object-level unlearning.

by read1 min views3 publishedOct 9, 2026

arXiv:2610.10859v1 Announce Type: new Abstract: Text-to-image diffusion models are increasingly distilled into few-step variants and being deployed to enable fast inference. However, their ability to generate harmful or undesired content poses significant safety risks. Data-driven unlearning methods suppress targeted generations by fine-tuning model weights using specialized unlearning objectives. Crucially, these objectives implicitly rely on multi-step denoising dynamics, an assumption that breaks down for few-step distilled (FSD) models, resulting in ineffective forgetting. Furthermore, performing unlearning on the non-distilled base model and subsequently re-distilling it to obtain an unlearned FSD model incurs substantial computational and time overhead, making it impractical in many settings. Hence, we address this limitation with a preference-driven unlearning framework that revisits Direct Preference Optimization (DPO) for diffusion models. We show that standard DPO and its unlearning derivatives, formulated around noise-prediction error, transfer poorly to FSD models due to their altered generation dynamics. To overcome this, we introduce a modified preference optimization formulation explicitly aligned with the few-step generation properties, enabling direct concept removal in FSD models while preserving few-step efficiency and maintaining strong retention of desirable (non-targeted) capabilities. We evaluate our framework primarily on identity and NSFW (nudity) removal tasks and also extend our method to object-level unlearning. Extensive experiments demonstrate consistent and effective forgetting, and strong retention performance, establishing our method as a practical and principled solution for unlearning in FSD models.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/enabling-preference-…] indexed:0 read:1min 2026-10-09 · —