3.1 Million Rollouts Later: No Agent Harness Wins Universally, and OpenEvolve Loses to Simpler Alternatives
Gupta, Lei, Lu, Anumanchipalli and Choshen (arXiv:2607.18235, July 20) decomposed popular automated-discovery frameworks including OpenEvolve and TTT-Discover into components, then ran 30 harness configurations across 12 model-problem pairs — over 3.1 million LLM rollouts with multi-trial statistical analysis. The headline: no fixed harness is reliably superior across model-problem pairs, and OpenEvolve variants generally underperform simpler alternatives. They argue harness choice is a hyperparameter needing per-problem tuning, and propose adaptive allocation that starts several harnesses, prunes weak runs early, and reallocates compute. Run pools and baseline distributions were released. For builders picking an agent scaffold, this is evidence the benchmark-winning harness may not be yours.
↳ Follow the thread