PreviewDiff introduces a training-free test-time search method that reconfigures diffusion sampling from scalar selection into multimodal critic-guided search over intermediate latents, addressing persistent failures in compositional fidelity such as object counting, attribute binding, spatial relations, and temporally grounded actions. Unlike Best-of-N sampling, which only selects among completed outputs, PreviewDiff intervenes at selected denoising checkpoints: it decodes a partial preview, queries a multimodal judge for both a score and natural-language critique, then branches over semantic prompt edits and locally re-noised latent continuations. These branches are scored and selectively rolled forward, allowing verifier compute to steer generation while the trajectory remains editable. Evaluated on image and video generation benchmarks, PreviewDiff consistently outperforms budget-matched Best-of-N and strong scalar-search baselines. Ablation studies reveal that earlier interventions and increased search width yield the largest gains, while deeper search and additional semantic variants provide complementary improvements. The results demonstrate that multimodal feedback functions most effectively not merely as a final verifier but as an active controller embedded within the denoising process itself.

Read original