We study length inflation in on‑policy distillation, where student responses become excessively long due to a termination‑token mismatch between base students and post‑trained teachers. Across Qwen3, Llama, and Gemma models, the stopping probability is placed on different EOS tokens even when their declared stopping sets are identical. This mismatch can suppress the student’s preferred termination and exhaust the generation budget.
Read original
huggingface/daily-papers