CROSS-CATEGORY: Three Independent Large-N Agent Evals Published in 48 Hours, All by Individuals, All Finding the Simple Configuration Wins
Between 2026-09-07 and 2026-09-08, three unrelated parties published high-sample-size adversarial evaluations of agents doing real work: Dan Luu's 26-condition, 160-runs-each testing study; Bottleneck Labs' seven-model, 72-hour autonomous-business run; and alvins82's 10 model/harness matrix on one Three.js task. None is a vendor benchmark, all publish per-condition raw numbers, and all three land on the same shape of result: the elaborate configuration (formal methods, more autonomy, the expensive model) loses to the plain one. This is the inverse of the vendor-publishes-eval-of-a-model-it-does-not-own format noted on 2026-09-05, and it produces the numbers a buyer can actually act on because the failure cases are itemized rather than aggregated.
↳ Follow the thread