Research
Audit of 150 Computer-Use Agent Trajectories Finds 15.3% of FAIL Verdicts Are Simply Wrong
Researchers audited 150 public failure-scored trajectories from five web, enterprise-workflow, and desktop-control computer-use benchmarks and found that 15.3% of FAIL verdicts were incorrect — 10.7% evaluator false negatives where a valid alternative solution was rejected, and 4.7% broken or stale tasks. For genuine failures, a three-tier diagnostic taxonomy shows verification/feedback and planning failures dominate execution and grounding errors, meaning a single scalar success rate hides the actual cause. Anyone comparing CUA products on published leaderboard numbers is reading a figure with roughly one in six negative results miscounted.
Source
↳ Follow the thread