Pattern: 'loop engineering' has a name, a benchmark, and a published industrial workflow in the same week
Three papers in four days converge on the outer loop as the unit of engineering rather than the prompt or the model: LoopArena (2608.28281) benchmarks a controller model separately from the coding agent it drives, Infobip's phased workflow (2608.30701) documents four phases with per-phase context strategies from production use, and the verification-surface study (2608.28795) holds the model fixed and varies only the agent's self-checking tools across 1,116 applications. The shift from last week's cluster is subtle but real: those papers argued the harness matters, these three are trying to measure and name specific loop decisions. If you maintain an outer loop, the vocabulary and the evaluation method now exist.
Source
↳ Follow the thread