{"slug": "past-prompt-adaptive-sampling-termination-for-efficient-diffusion-model", "title": "PAST: Prompt-Adaptive Sampling Termination for Efficient Diffusion Model", "summary": "Researchers propose PAST, a prompt-adaptive sampling termination method that improves reinforcement learning fine-tuning for diffusion models by up to 66.7% in computational efficiency and up to 29.5% in preference optimization quality, according to an arXiv paper (2608.06794v1). PAST provides differentiated rewards, adaptively regulates training episode length based on denoising progress and prompt difficulty, and uses a dual adaptive coordination mechanism to balance exploration and convergence.", "body_md": "arXiv:2608.06794v1 Announce Type: new\nAbstract: While diffusion models have made significant progress in text-to-image tasks, they still exhibit limitations when directly optimizing downstream objectives. Although Reinforcement Learning (RL) enables targeted optimization, existing methods are generally constrained by low-efficiency fine-tuning and sparse rewards. To address these challenges, we propose PAST, which provides differentiated rewards while adaptively regulating training episode length by jointly perceiving denoising progress and prompt difficulty. Specifically, we design an intrinsic reward paradigm to compensate for sparse extrinsic rewards and guide the model to explore paths that diverge more efficiently from noise patterns. We further provide theoretical justification for intrinsic rewards. Then, PAST dynamically monitors denoising completion and semantic alignment between image structures and prompt semantics. When both metrics satisfy generation requirements, the system adaptively terminates training. This enables appropriate allocation of episode lengths based on prompt difficulty and the current generation process. Finally, based on the predicted residual noise level, we establish a dual adaptive coordination mechanism. Specifically, it not only balances the extrinsic and intrinsic rewards but also balances the exploration and convergence. Experimental results demonstrate that PAST enhances computational efficiency of existing RL fine-tuning methods by up to 66.7%, while improving preference optimization quality by up to 29.5% through its dual adaptive regulation mechanism.", "url": "https://wpnews.pro/news/past-prompt-adaptive-sampling-termination-for-efficient-diffusion-model", "canonical_source": "https://arxiv.org/abs/2608.06794", "published_at": "2026-08-10 04:00:00+00:00", "updated_at": "2026-08-10 04:14:59.409322+00:00", "lang": "en", "topics": ["machine-learning", "artificial-intelligence"], "entities": ["arXiv", "PAST"], "alternates": {"html": "https://wpnews.pro/news/past-prompt-adaptive-sampling-termination-for-efficient-diffusion-model", "markdown": "https://wpnews.pro/news/past-prompt-adaptive-sampling-termination-for-efficient-diffusion-model.md", "text": "https://wpnews.pro/news/past-prompt-adaptive-sampling-termination-for-efficient-diffusion-model.txt", "jsonld": "https://wpnews.pro/news/past-prompt-adaptive-sampling-termination-for-efficient-diffusion-model.jsonld"}}