Sources
OSReward Finds VLM Judges Systematically Score Agent Failures as Successes, and Ships a 9B/35B Replacement at 30–60x Lower Cost
With no way to hand-verify computer-use agent trajectories at scale, the field has defaulted to VLM-as-judge; this benchmark (arXiv 2607.28609, revised Aug 6, 22+ authors led by Qiushi Sun) builds human-verified ground truth for those judgments and finds even state-of-the-art models fall short, with a consistent bias toward misclassifying failures as successes. The authors release OS-Shepherd at 9B and 35B, trained on a 100K corpus, claiming reliable judging at 30–60x lower cost than frontier models. Note this is a revision of a July 30 submission rather than a brand-new paper, but the finding is directly load-bearing for anyone whose agent eval pipeline trusts a model grader.
↳ Follow the thread