The paper introduces a unified, controlled multi‑turn environment that enables systematic study of long‑horizon planning in foundation model agents across distinct training stages. By replacing opaque internet data with a precise, controllable setting, the authors aim to clarify how planning ability is acquired, shaped, and integrated. This framework supports analysis from pre‑training through post‑training via single‑ and multi‑teacher on‑policy agentic distillation.

Read original