The authors reinterpret on‑policy distillation (OPD) as a reinforcement‑learning problem, showing that the reverse‑KL objective used in OPD is equivalent to a KL‑regularized policy‑optimization objective. Building on this equivalence they propose Least‑Square Policy Distillation (LSPD), which injects two RL mechanisms—optimistic exploration and off‑policy reuse of previously gathered trajectories—into the distillation loop. LSPD therefore maintains policy diversity through exploration while improving rollout efficiency by repeatedly learning from stored data. A theoretical analysis links LSPD to optimistic value‑based methods and proves that its idealized form attains a Õ(log K) regret bound under online exploration. Empirically, LSPD is evaluated on six mathematical‑reasoning benchmarks with various teacher‑student model pairs; it outperforms prior distillation baselines by an average of +1.59 points in Avg@16. Pass@k measurements up to k=64 reveal that LSPD’s advantage grows with k, indicating better preservation of diverse reasoning policies. Moreover, a fully off‑policy variant of LSPD matches the performance of vanilla OPD while consuming only the first 25 % of rollout batches, demonstrating substantial sample‑efficiency gains.

Read original