Iris is an open-weight web research agent trained using a combined SFT-RL climbing approach to handle complex multi-hop search queries that typically cause traditional agents to fail. The training methodology addresses the common breakdown in agent performance when tasks require tracing chains of information across multiple steps. By integrating supervised fine-tuning with reinforcement learning, Iris learns to maintain context and reason through intricate web research tasks more effectively.
Read original
dev.to