Fetching from the wire…
Top 5 · 2026-08-11 · source-backed
Best-in-class computer-use models scored 42% on OSWorld-Verified in early 2025. Today the leader (Claude Fable 5) scores 85%. The human tester baseline is roughly 72%. a16z published the aggregation on August 10, pulling from production interviews and llm-stats leaderboard data.
Crossing the human baseline is the eye-catching part, but the cost math is what changes decisions. Pure screenshot-loop operation runs $6–8/hour of inference, with the full range spanning $3–15 depending on how the harness is designed. Offshore BPO runs about $10/hour fully loaded. US back-office labor runs $30–45/hour. So the agent is roughly break-even against offshore and carries a 70–80% gross margin against domestic.
That's a very specific place to be. Not "cheaper than everything," which would have triggered instant commoditization. Not "still too expensive," which would have kept this in demo-land. Break-even against the cheapest human option and profitable against the expensive one, right now, at today's prices, which are falling.
The strategic read from the piece is the one builders should internalize: raw UI navigation has commoditized down into the model layer. If your product's value proposition was "we can click through a legacy web app reliably," the model does that now for $7/hour. The advantage moved up the stack to context (knowing which app, which account, which state), permissions (what the agent may touch), process knowledge (what the workflow actually is versus what the SOP says), validation (did it work), escalation (when to get a human), and run-caching (don't pay twice for the same trajectory).
One founder quote in the piece deserves attention for its precision: "the models weren't good enough to use in production on their own until Opus 4.6 in February 2026." That's six months ago. An entire product category became viable half a year ago, which means most of the durable companies in it haven't been founded yet.
This converges with Ouroboros reporting 90.69% on OSWorld-Verified from a completely different direction. Two independent sources putting computer-use well past the human baseline in the same week is the kind of agreement that's hard to dismiss as leaderboard gaming.
What I'd do with this: stop building the clicking. Start building the wrapper. If you have a workflow automation product, your roadmap for the next two quarters is process capture, permission scoping, and failure escalation, not better selectors. And if you're evaluating whether to buy or build here, note that $6–8/hour is a real operating cost that scales linearly, not a fixed-cost software play. The unit economics look more like staffing than like SaaS.
Each link below shares sources, entities, or timing with this story.
Claude Fable built by Anthropic / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Claude Fable built by Anthropic); both cover Opus, Start; reported by the same outlet (arxiv.org).
Claude Code uses Opus / Shared entities / Earlier coverage / Tension
Linked by a graph relationship (Claude Code uses Opus); both cover Opus, OSWorld, Verified; earlier Opus coverage from 2026-02-17.
Claude Fable built by Anthropic / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude Fable built by Anthropic); both cover Opus, Verified; overlapping topics (cost, model).
Claude Code uses Opus / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude Code uses Opus); both cover August, Opus; overlapping topics (against, agent).
Claude Fable built by Anthropic / Shared entity: Verified / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Claude Fable built by Anthropic); both cover Verified; reported by the same outlet (arxiv.org).
Cursor uses Opus / Shared entity: Verified / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Cursor uses Opus); both cover Verified; reported by the same outlet (arxiv.org).
Cursor uses Opus / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Cursor uses Opus); both cover Opus, SaaS; overlapping topics (agent, cost).
Claude Fable uses CUDA / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Claude Fable uses CUDA); both cover Claude Fable, Opus; overlapping topics (baseline, model).