Agents
A manager-worker scaffold beats a larger model on cost, but the gain is null or negative for a third of models tested
A 27 August study isolates the manager-worker scaffold over a shared filesystem workspace against the same model answering in a single pass, on the 100 latest hard LiveCodeBench problems across nine models from 9B to about 2.8T parameters. The scaffold gives large significant gains for some (Kimi-K3 +30.4, Qwen3.8-27B +23.4, GPT-5.6-Luna +10.6) and is null to negative for others (Qwen3.6-35B -1 to -9 with reasoning off). Running a manager roughly triples the token bill but buys accuracy more cheaply than moving to a larger model, and Opus-5 with the manager posts the study's top score at 91% in one pass.
Source
↳ Follow the thread