Researchers investigated on-policy distillation (OPD) for language models at the data-minimal limit, discovering that training on a single query enables continued improvement for hundreds of steps while recovering most