Tools
Alibaba's Accio Team Publishes RealReplicaBench: 107 Long-Horizon Agent Tasks in Stateful Replicas of Real Services
Accio-Lab/RealReplicaBench appeared 2026-08-02 (162 stars, release v1.3.1) from the Accio team at Alibaba International, benchmarking long-horizon agents against high-fidelity, stateful, reproducible replicas of real online commerce services rather than mocks or static transcripts. It ships 107 tasks, a reproducibility contract, a live leaderboard, and reference results run through the OpenClaw harness. The stateful-replica approach targets the specific failure mode static benchmarks can't measure: whether an agent's earlier actions leave the world in a state its later actions can still work in.
Source
↳ Follow the thread