T1 is a 122B-parameter Mixture-of-Experts model trained via reinforcement learning to handle long-horizon terminal tasks such as coding and scientific discovery. It operates a real shell in a cloud sandbox, supporting up to 300+ tool-call turns per task, and is rewarded by executing each task's own verifier. The training recipe employs an aggressively warm-started approach to stabilize actor-critic learning.

→ View original source