OSReward introduces a standardized evaluation framework for cross-platform computer-use reward models, addressing scalability challenges in verifying CUA trajectories through vision-language models (VLMs) as human annotators prove impractical. The approach focuses on automating task instruction verification to enhance data curation and reinforcement learning for advancing CUAs. Developed by researchers from multiple institutions, it aims to standardize assessment criteria across diverse platforms. Read original