The paper introduces SparsePR, a training-free block-sparse attention mechanism for video transformers that addresses limitations in existing sparse attention by analyzing how partition geometry affects pooled support and residual predictability. This method aims to accelerate video generation and world models without requiring additional training.

→ View original source