Research
LoopArena Scores Models as Loop Controllers Directing a Separate Coding Agent, Best Strict Success 24.69%
LoopArena (arXiv 2608.28281, submitted 2026-08-28, cs.AI) benchmarks 'Loop Engineering' directly: a Controller model receives a structured summary after each coding round and tells a separate fixed Worker agent what to do, verify, or when to stop, isolating loop guidance from the coding agent's own ability. On full tasks the best observed Strict Success Rate is 24.69%, and across Controllers the paired reduction in estimated inference cost averages 64.4%. The cheap Type II setting reproduces the full-task ordering at Spearman rho=0.9747, so builders can rank their own loop controllers without paying for full end-to-end runs; data and code are at github.com/AMAP-ML/LoopArena.
↳ Follow the thread