State-of-the-art video diffusion models produce visually impressive yet physically implausible content. An interpretability study shows that these violations arise from attention‑driven motion planning during early denoising stages. Improving internal attention mechanisms, rather than relying only on external priors, may enforce physical consistency.
Read original
huggingface/daily-papers