Fetching from the wire…
Public story · 2026-08-07 · high
The paper's human-verified benchmark shows the bias holds even in frontier models, and its replacement judge claims a 60x cost cut.
Why now: The paper is a revision of a July 30 submission that surfaced in the August 7 coverage, and the bias it documents doesn't become less relevant with age.
Vision-language judges score failed computer-use runs as successes far more often than they should, per a new benchmark called OSReward.
That's a problem for anyone grading agent runs with a model instead of a person, since a biased judge inflates every success rate above it. Hand-verifying trajectories doesn't scale, so most eval setups already lean on a model grader.
The benchmark tests state-of-the-art vision-language models against human-labeled ground truth on computer-use trajectories. It finds a consistent skew: judges misclassify failures as successes more often than the reverse. The authors respond with OS-Shepherd, a purpose-built judge released at 9B and 35B parameters, trained on a 100,000-example corpus of human-verified trajectories. They claim it matches frontier-model judging accuracy at 30 to 60 times lower cost.
Yes, this is a revision of a July 30 submission, not new research. But the bias it documents doesn't expire, and most teams still haven't checked whether their own judge model has it.
If your agent eval grades trajectories with an LLM and skips spot-checking a sample by hand, your reported success rate is probably inflated. Worth watching whether teams building computer-use agents start swapping in something like OS-Shepherd, or keep trusting judges the paper shows are biased toward false positives.
Each link below shares sources, entities, or timing with this story.
Shared entity: July / Same source domain / Shared topic / Earlier coverage / Tension
Both cover July; reported by the same outlet (arxiv.org); overlapping topics (cost, model).
Both cover July; reported by the same outlet (arxiv.org); overlapping topics (agent, model).
Both cover July; reported by the same outlet (arxiv.org); overlapping topics (agent, model).
Both cover July; reported by the same outlet (arxiv.org); overlapping topics (agent, model).
Shared entity: VLM / Same source domain / Shared topic / Earlier coverage / Tension
Both cover VLM; reported by the same outlet (arxiv.org); overlapping topics (directly, model).
Shared entity: VLM / Same source domain / Shared topic / Earlier coverage
Both cover VLM; reported by the same outlet (arxiv.org); overlapping topics (agent, cost, eval).
Shared entity: July / Shared topic / Earlier coverage
Both cover July; overlapping topics (agent, claiming, cost, eval, model); earlier July coverage from 2026-07-17.
Shared entities / Shared topic / Earlier coverage
Both cover July, Shepherd; overlapping topics (agent, doesn); earlier July coverage from 2026-07-16.