Fetching from the wire…
Agents2026-08-02 · source-backed
Researchers reviewed public failure-scored trajectories from five web, enterprise-workflow, and desktop-control benchmarks: 10.7% were evaluator false negatives rejecting valid alternative solutions, 4.7% were broken or stale tasks. For the genuine failures, verification/feedback and planning errors dominate execution and grounding. Anyone comparing CUA products on leaderboard numbers is reading a figure with one in six negative results miscounted.
Each link below shares sources, entities, or timing with this story.
Shared entity: Anyone / Same source domain / Shared topic / Earlier coverage / Tension
Both cover Anyone; reported by the same outlet (arxiv.org); overlapping topics (agent, anyone, benchmark, evaluator).
Shared entity: Anyone / Same source domain / Shared topic / Earlier coverage
Both cover Anyone; reported by the same outlet (arxiv.org); overlapping topics (agent, anyone, benchmark, broken).
Shared entity: Researchers / Same source domain / Shared topic / Earlier coverage / Tension
Both cover Researchers; reported by the same outlet (arxiv.org); overlapping topics (agent, broken).
Shared entity: Anyone / Same source domain / Shared topic / Earlier coverage
Both cover Anyone; reported by the same outlet (arxiv.org); overlapping topics (anyone, audit, benchmark).
Both cover Anyone; reported by the same outlet (arxiv.org); overlapping topics (agent, anyone, benchmark).
Shared entity: Audit / Same source domain / Shared topic / Earlier coverage
Both cover Audit; reported by the same outlet (arxiv.org); overlapping topics (agent, audit, execution).
Shared entity: Anyone / Shared topic / Earlier coverage / Tension
Both cover Anyone; overlapping topics (anyone, benchmark, comparing); earlier Anyone coverage from 2026-07-28.
Shared entity: Audit / Shared topic / Earlier coverage / Tension
Both cover Audit; overlapping topics (agent, audit, evaluator); earlier Audit coverage from 2026-06-28.