Qwen-Image-2.0-RL Technical Report
Qwen-Image-2.0-RL introduces a post-training pipeline utilizing reinforcement learning from human feedback (RLHF) and on-policy distillation (OPD) to enhance the Qwen-Image-2.0 diffusion model. The framework employs task…
→ View original source