MerchantBench Runs Agents Through 365 Simulated Days of E-Commerce: the Best LLM Reached 27.3% of Human Final Net Assets and Stayed Engaged Only 10.6–66.1% of the Time
Topping Hugging Face Daily Papers on August 5 with 74 upvotes, MerchantBench (arXiv 2607.28956, posted July 31, 2026) grounds a year-long e-commerce simulation in 98,843 real product records, forcing agents to coordinate sourcing, pricing, cash management and order handling under mixed-latency feedback. Humans finished at ~217,610 RMB in net assets; the strongest model, GPT-5.6 Sol, reached 40,890 RMB under ReAct and 52,930 RMB under Hermes — 27.3% of the human mean, with the Hermes scaffold worth 53.3% more than ReAct on average. The failure taxonomy is the useful part for anyone shipping long-running agents: "operational coherence" decay (activity trails off) and "strategic coherence" breakdown (goal drift, stops updating on evidence), with humans at 100% sustained engagement versus 10.6–66.1% for LLMs.
↳ Follow the thread