cd /news/machine-learning/past-prompt-adaptive-sampling-termin… · home topics machine-learning article
[ARTICLE · art-89863] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=· neutral

PAST: Prompt-Adaptive Sampling Termination for Efficient Diffusion Model

Researchers propose PAST, a prompt-adaptive sampling termination method that improves reinforcement learning fine-tuning for diffusion models by up to 66.7% in computational efficiency and up to 29.5% in preference optimization quality, according to an arXiv paper (2608.06794v1). PAST provides differentiated rewards, adaptively regulates training episode length based on denoising progress and prompt difficulty, and uses a dual adaptive coordination mechanism to balance exploration and convergence.

read1 min views1 publishedAug 10, 2026

arXiv:2608.06794v1 Announce Type: new Abstract: While diffusion models have made significant progress in text-to-image tasks, they still exhibit limitations when directly optimizing downstream objectives. Although Reinforcement Learning (RL) enables targeted optimization, existing methods are generally constrained by low-efficiency fine-tuning and sparse rewards. To address these challenges, we propose PAST, which provides differentiated rewards while adaptively regulating training episode length by jointly perceiving denoising progress and prompt difficulty. Specifically, we design an intrinsic reward paradigm to compensate for sparse extrinsic rewards and guide the model to explore paths that diverge more efficiently from noise patterns. We further provide theoretical justification for intrinsic rewards. Then, PAST dynamically monitors denoising completion and semantic alignment between image structures and prompt semantics. When both metrics satisfy generation requirements, the system adaptively terminates training. This enables appropriate allocation of episode lengths based on prompt difficulty and the current generation process. Finally, based on the predicted residual noise level, we establish a dual adaptive coordination mechanism. Specifically, it not only balances the extrinsic and intrinsic rewards but also balances the exploration and convergence. Experimental results demonstrate that PAST enhances computational efficiency of existing RL fine-tuning methods by up to 66.7%, while improving preference optimization quality by up to 29.5% through its dual adaptive regulation mechanism.

── more in #machine-learning 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/past-prompt-adaptive…] indexed:0 read:1min 2026-08-10 ·