Fetching from the wire…
Public story · 2026-08-28 · high
The gains split across nine models sized 9B to 2.8 trillion parameters, each tested with a manager coordinating worker agents on coding tasks.
Why now: The paper posted August 27, testing each model against the 100 latest hard problems on LiveCodeBench.
A manager-worker AI coding scaffold beat a single model working alone on most systems tested, but failed for a third of them, per a paper posted August 27.
For teams building multi-agent coding tools, that split decides whether a manager agent is worth the extra token spend. Running one triples the bill, and for a third of the nine models in the study, it bought nothing or made results worse.
The manager-worker LiveCodeBench study ran nine models from 9B to about 2.8 trillion parameters against the 100 latest hard problems on LiveCodeBench. Each model got a shared filesystem workspace to coordinate through.
Kimi-K3 gained 30.4 points with a manager coordinating worker agents. Qwen3.8-27B gained 23.4 points, and GPT-5.6-Luna gained 10.6. Qwen3.6-35B lost 1 to 9 points when run with reasoning turned off.
The token cost triples when a manager runs the show. But the paper's comparison is against moving to a larger model, not against asking one model to answer alone. Opus-5 with the manager posted the study's top score, 91% in one pass.
Each link below shares sources, entities, or timing with this story.
Same source domain / Semantically similar
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.77).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.76).
Same source
Cite the same source (arXiv).
Same source domain / Semantically similar
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.74).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.74).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.73).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.73).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.73).