Rufus‑Air presents an open, reproducible post‑training recipe applied to the GLM‑4.5‑Air‑Base model (106B‑A12B parameters). The recipe consists of eight serial stages: supervised fine‑tuning (SFT), reasoning reinforcement learning (RL), coding RL, instruction‑following RL, a general‑purpose agent stage, a coding‑agent stage, a search‑agent stage, and finally RLHF. Each stage moves from basic to more advanced capabilities and shifts reward signals from hard, verifiable metrics to softer, judge‑based evaluations. The authors detail the datasets, reward functions, infrastructure, stage ordering, and per‑stage results required to replicate the pipeline. Training relies exclusively on openly available components and public data, using them as‑released without additional human annotation or an internal distillation teacher. Key observations include that high‑quality, diverse SFT establishes a strong capability baseline, that difficulty filtering keeps RL prompts within a productive learning window, that reward reliability guides stage ordering, and that engineering and infrastructure decisions are integral to the recipe. Rufus‑Air outperforms the official GLM‑4.5‑Air post‑trained release and remains competitive with other open models of comparable size.
Read original
huggingface/daily-papers