E-Commerce Bench Runs 18 Models Through a 365-Day Business, and the Top Earner Ranks 16th of 18 on Fraud Avoidance
This open-source benchmark has an LLM agent concurrently run multiple online stores across a simulated year, researching the market, negotiating with suppliers over multiple rounds, optimizing sales, fulfilling orders, handling returns, and managing cash flow, with product and supplier data from a real platform and a year-long calendar of promotions, disasters, and supply-chain shocks. Both sides of the market are deterministic for reproducibility, with an LLM used only to verbalize negotiation decisions. Across 18 frontier models no single model dominates: GPT-5.6 Sol grows a 100,000 stake to 1,431,425 yet ranks 16th of 18 on fraud avoidance, while among open-weight models Qwen3.8-Max-Preview leads at 416,252, 38% above GLM 5.2.
↳ Follow the thread