This survey examines post‑training and alignment strategies for video generation models, framing post‑training as a unified approach that adapts pretrained generators without full retraining. The authors distinguish implicit alignment, where alignment signals influence model behavior indirectly, from explicit alignment, which directly enforces such signals during training or inference. Building on this dichotomy, they categorize existing techniques into four groups: supervised fine‑tuning, which uses labeled video‑text pairs to adjust model weights; self‑training and distillation methods, which generate pseudo‑labels or transfer knowledge from larger models; preference‑ and reward‑based methods, which optimize models using human or automated feedback signals; and inference‑time methods, which modify sampling or decoding processes to steer output toward desired attributes. The paper also surveys the datasets, benchmarks, and evaluation metrics commonly employed to assess temporal coherence, motion‑appearance fidelity, and safety constraints in generated videos. It highlights open challenges such as designing scalable reward functions that capture long‑horizon dynamics, maintaining stability while preserving expressive motion, mitigating error accumulation over extended sequences, and ensuring safety‑aware generation amid multi‑objective trade‑offs. By consolidating these methodological advances and practical considerations, the survey provides a structured foundation for developing controllable and reliable video generation systems.

Read original