Researchers propose a person-aligned user simulation framework for evaluating role-playing agents (RPAs) in interactive, multi-turn conversations. Unlike traditional benchmarks that rely on fixed dialogue histories and static evaluation rubrics, this approach aligns simulated users with real user characteristics to enable more dynamic and realistic assessment. The method aims to improve the reliability of RPA evaluations by better reflecting actual user interactions and expectations. This work addresses the need for evaluation systems that can adapt to diverse user behaviors rather than following a rigid, one-size-fits-all approach.
Read original
huggingface/daily-papers