Research
OSWorld-Pro Grades Computer-Use Agents on 2,800 Subgoals, and Claude Opus 5 Drops From 83.4% to 75.7%
OSWorld-Pro (arXiv 2609.24890, 21 Sep 2026) replaces end-state-only computer-use evaluation with process-based scoring: 300+ tasks decomposed into over 2,800 sequentially dependent subgoals, grounded in more than 67,000 human annotations and judged by human-aligned LLM judges. Top performer Claude Opus 5 reaches only 75.7% versus 83.4% on the original OSWorld, and the subgoal traces separate distinct failure modes such as subgoal-irrelevant actions from click-based grounding mistakes. For builders, that distinction is the point: keyboard-input errors and GUI-click errors need different mitigations, which final-deliverable scoring cannot tell apart.
Source
↳ Follow the thread