Six months of production LLM trading agents: the operating layer beat strategy text, and neither fleet had an edge
A population-scale record across two live systems (3,505 user-funded vaults trading real ETH for 21 days, plus a 500-599 agent fleet on Hyperliquid perpetuals over three months) covering 7.5M model invocations, ~300K onchain actions and 14,596 fills found the harness dominates the prompt: a risk slider explains leverage at +0.425 per level, agent fixed effects absorb 60% of variance, and a leaderboard render boundary causally routed selection (regression discontinuity 1.75x at the top-3 cut). Sizing was volatility-blind, with median leverage 5.0x in every volatility sextile and one slider cell holding 11% of the book but 62% of liquidations. 43.2% of positions saw at least +300 bps of favorable excursion within 24h yet 49.3% of those closed negative, a mechanical bracket recovers +39.0 bps per position, and a paired replay of frontier models on 416 captured production scenarios found decision quality statistically indistinguishable while choice stability differed sharply by model family.
↳ Follow the thread