This paper investigates dialogue game agents using the LM Playschool Challenge with a 2B open-weight model, identifying that failures stem not only from broad knowledge gaps but also from local decision errors such as repeated guesses, malformed actions, and ignored feedback. The authors propose a diagnosis-guided post-training recipe—acquire, repair, and preserve—targeting these specific failure modes. The approach addresses the challenge of maintaining state across turns and selecting valid actions under evolving constraints in interactive dialogue settings.

Read original