AV‑GRPO introduces a modality‑anchored online diffusion reinforcement‑learning framework for joint audio‑video generation that tackles heterogeneous rewards, divergent modality dynamics, and synchronization evaluation bias. The method comprises three modules: (1) modality‑anchored rollouts that generate separate audio and video trajectories to disentangle learning signals and keep difficulty stable; (2) trajectory‑locked frozen‑tower optimization, where one modality tower is frozen while the other is updated, reducing computational cost and reassigning credit more cleanly; and (3) adaptive objectives and perturbation strengths tuned to each modality’s intrinsic dynamics, converting coupled multimodal preference learning into unimodal subproblems for precise reward attribution and better cross‑modal sync. Complementarily, the authors release 5DAV, a decoupled dataset that varies samples along five controllable dimensions (e.g., audio‑visual temporal offset, semantic complexity, motion intensity, texture detail, and background noise) to enable systematic training. Experiments on JavisBench and VABench show AV‑GRPO surpasses the LTX‑2.3 baseline in generation quality, semantic alignment, and cross‑modal synchronization under both LoRA adapters and full fine‑tuning, with ablation studies confirming the contribution of each module. Code and data are publicly available.
Read original
huggingface/daily-papers