SEED (SElf-Evolving On-Policy Distillation) addresses the supervision gap in outcome-based reinforcement learning for large language model agents by enabling self-evolving, on-policy distillation. The method improves sparse trajectory-level rewards into token-level policy learning guidance, enhancing multi-turn interaction and tool use. Authors propose SEED as a practical optimization paradigm for long-horizon tasks.