a16z Puts Numbers on Computer-Use Agents: OSWorld-Verified Went 42% → 85% in a Year, Past the ~72% Human Baseline, at $6–8/Hour of Inference
An a16z research piece published Aug 10 aggregates production interviews and llm-stats leaderboard data: the best computer-using model scored 42% on OSWorld-Verified in early 2025 versus 85% today (Claude Fable 5 leading), against a ~72% human tester baseline. Cost lands at roughly $6–8/hour of inference in pure screenshot-loop mode ($3–15 depending on harness design), which is about break-even against offshore BPO at ~$10/hour fully loaded and a 70–80% gross margin against US back-office labor at $30–45/hour. The builder-relevant conclusion is that raw UI navigation has commoditized to the model layer and the durable advantage moved up the stack — context, permissions, process knowledge, validation, escalation, and run-caching; as one founder put it, 'the models weren't good enough to use in production on their own until Opus 4.6 in February 2026.'
↳ Follow the thread