OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
OSReward is introduced, a realistic, high-quality benchmark that evaluates VLM judges on CUA trajectories, and an open corpus of reasoning-annotated trajectory judgments for the CUA community, to close the gap in reliable CUA reward at scale.