Fetching from the wire…
Public story · 2026-07-31 · high
Four Opus 4.6 agents coordinating through the system beat a solo agent on the newer Opus 4.8 model, 62.1% to 57.2%, per the July 30 paper.
Why now: The paper posted to arXiv on July 30.
AgentRadio lets coding agents message each other mid-task, and four coordinated agents resolved 62.1% of codebase questions versus 32.3% for one agent working alone, per a July 30 arXiv paper. That's a nearly 30-point jump in benchmark accuracy. It comes from a coordination layer teams could add to agents they already run, not a newer model.
The system adds three primitives: threads, messages, and waiting for a mention. The wait runs as a background task, so an agent keeps working in the foreground instead of stopping to check in.
The benchmark is SWE-Atlas QnA, a set of long-horizon questions about production codebases. The four-agent group split the work using a five-phase division-of-labor protocol, and the code is released under Coral-Protocol.
That 62.1% score also beats a single Claude Code agent running the newer Opus 4.8 model, which answers 57.2% of the same questions.
The accuracy gap between solo and coordinated agents widens as tasks get harder, per the paper. The authors read that as mid-course correction, agents catching and fixing each other's wrong turns, not four agents splitting a workload.
A related paper, SkillRise, takes a different route to the same goal. It has a single agent alternate between solving tasks and rewriting its own skill document. No teammates required, one policy improving itself instead of four agents talking to each other.
Each link below shares sources, entities, or timing with this story.
Same source domain / Semantically similar
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.76).
Same source
Cite the same source (arXiv).
Same source domain / Semantically similar
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.75).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.75).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.74).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.74).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.73).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.73).