Tools
RealReplicaBench Puts Claude Opus 5 on Top at 61.7% Across 107 Long-Horizon Tasks in Cloned Commerce Services
Accio-org/RealReplicaBench (created 2026-08-02, 1,038 stars in 5 days) benchmarks agents against high-fidelity stateful replicas of eight commerce and logistics platforms — Alibaba-style publishing, FreightOS-style freight, Shopify-style storefront admin — over 107 tasks (53 CLI, 28 browser, 16 file, 10 API/MCP). Twelve models were run on two harnesses: Claude Opus 5 leads at 66/107 (61.7%) on the Accio harness and 60/107 (56.1%) on OpenClaw, ahead of Opus 4.8 (55.1%/51.4%) and GPT-5.6 Sol (49.5%). The harness swap moves scores by 5+ points on identical models — a reminder that agent scaffolding is a measured variable, not a footnote.
Source
↳ Follow the thread