On-policy diffusion distillation (OPD) extends velocity matching to classifier-free guidance (CFG)-composed predictions, but its branch-level behavior remains under-identified. The study reveals that directly matching teacher and student guided velocities may lead to suboptimal performance due to under-identification at the branch level. This challenges the default assumption in modern diffusion systems and suggests a need for rethinking CFG integration in OPD.
Read original
huggingface/daily-papers