This paper introduces ActObs, a supervised fine-tuning method that extends standard practice by applying loss not only to agent action tokens but also to environment observation tokens present in trajectory data. While deployed agents never generate observations, the authors find that supervising them during SFT yields better initialization for subsequent reinforcement learning. The approach challenges the conventional assumption that observations should serve only as context rather than prediction targets.
Read original
huggingface/daily-papers