On-policy distillation (OPD) is analyzed as an exploration catalyst in LLM post-training, guiding students through token-level reasoning without increasing capability ceilings. The study emphasizes prompt diversity over sampling quantity for effective learning. Pathologies and regulatory mechanisms within OPD are systematically examined to enhance training dynamics understanding. Read original