The paper investigates an asymmetry in multimodal diffusion transformers where companion modalities such as 3D body motion or audio form strong attentional links to video tokens, but the reverse links that allow those modalities to constrain video generation remain weak. To quantify this, the authors compute correspondence distributions over video tokens for both video‑to‑modality and modality‑to‑video directions and define their disagreement as the reciprocal correspondence gap. They propose RecCAR (Reciprocal Cross‑modal Attention Regularization), a KL‑divergence regularizer that treats the well‑established video‑to‑modality distribution as a fixed reference and pulls the weaker modality‑to‑video distribution toward it. Experiments on joint video‑motion and video‑audio generation show that RecCAR raises the Human Anatomy score—a metric of pose fidelity—from 0.69 to 0.75 and lowers audio‑video desynchronization, measured as the average temporal misalignment error, from 0.804 to 0.752. These improvements indicate that enforcing reciprocal attention balance yields more coherent and anatomically plausible video outputs while preserving the generative quality of the model.

Read original