Sources
CommerceAgentBench scores agents on what they changed, not what they said, and the best passes 65 of 107 tasks
Alibaba International's Accio team open-sourced a 107-task benchmark (53 CLI, 28 browser, 16 file, 10 API/MCP) running against fourteen offline replicas of real business software in a fresh container per task, with verifiers inspecting mock-service state rather than the transcript. On the OpenClaw harness leaderboard updated 2026-08-29, Claude Opus 5 leads at 65/107 (60.7%) using 52.5 steps, 7.8 minutes and 2.05M tokens per task, ahead of Opus 4.8 at 56/107 and Qwen 3.8 Max and DeepSeek V4 Pro tied at 53/107. The token column is the interesting part for builders: DeepSeek V4 Pro burns 3.58M tokens to reach the same pass rate GPT-5.6 Sol nearly matches on 1.15M.
↳ Follow the thread