cd /news/artificial-intelligence/previewdiff-multimodal-critic-guided… · home › topics › artificial-intelligence › article
[ARTICLE · art-142261] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents

PreviewDiff, a training-free test-time search method described in arXiv paper 2609.36199v1, turns diffusion sampling from scalar search into a multimodal critic-guided search over intermediate latents, decoding partial previews at selected denoising checkpoints so a multimodal judge can score and critique them before generation fails. Across image and video generation benchmarks, PreviewDiff consistently improves over budget-matched Best-of-N selection and strong scalar-search baselines, with ablations showing earlier interventions and increased search width provide the largest gains. The authors report that multimodal feedback is most useful not only as a final verifier but as an active controller inside the denoising process.

by read1 min views1 publishedSep 30, 2026

arXiv:2609.36199v1 Announce Type: new Abstract: Diffusion models can produce striking images and videos, but they still struggle with the compositional details that make a generation faithful to a prompt, such as object counts, attribute binding, spatial relations, and temporally grounded actions. A common way to improve prompt satisfaction is to spend more compute at test time through Best-of-N sampling, but final-sample selection is fixed. Best-of-N can only choose among completed outputs and cannot repair a promising trajectory before it fails. We introduce PreviewDiff, a training-free test-time search method that turns diffusion sampling from scalar search into a multimodal critic-guided search over intermediate latents. At selected denoising checkpoints, PreviewDiff decodes a partial preview, asks a multimodal judge to score and critique it, and uses the resulting natural-language feedback to branch over semantic prompt edits and locally re-noised latent continuations. These branches are then scored and selectively rolled forward, allowing verifier compute to guide generation while the sample is still editable. Across image and video generation benchmarks, PreviewDiff consistently improves over budget-matched Best-of-N selection and strong scalar-search baselines. Ablations show that earlier interventions and increased search width provide the largest gains, while deeper search and additional semantic variants offer complementary improvements. PreviewDiff demonstrates that multimodal feedback is most useful not only as a final verifier, but as an active controller inside the denoising process.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @previewdiff 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/previewdiff-multimod…] indexed:0 read:1min 2026-09-30 · —