Fetching from the wire…
Agents2026-08-27 · source-backed
The authors separated multi-agent reasoning into candidate generation, peer communication and terminal selection, held two fixed to isolate the third, and replayed 81,390 fixed candidate pools drawn from 16,278 questions across five benchmarks (arXiv 2608.25937). A correct answer is frequently already in the pool while the system converges on a wrong one, a failure they call memetic drift. Judge reliability turns out not to be a fixed model trait but to vary with task, generator, and how rare the correct answer is. Combining answer frequency with the judge's verdict, changing only the selection rule, lifts accuracy.
Each link below shares sources, entities, or timing with this story.
Shared entity: Combining / Same source domain / Shared topic / Earlier coverage
Both cover Combining; reported by the same outlet (arxiv.org); overlapping topics (benchmark, multi-agent).
Shared entity: Judge / Same source domain / Earlier coverage / Tension
Both cover Judge; reported by the same outlet (arxiv.org); earlier Judge coverage from 2026-08-10.
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (accuracy, answer, author); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (accuracy, benchmark, call); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (answer, question, wrong); pushes against this story (against).
Same source domain / Shared topic / Downstream implication
Reported by the same outlet (arxiv.org); overlapping topics (judge, multi-agent, system); traces where this leads (downstream).
Same source domain / Shared topic
Reported by the same outlet (arxiv.org); overlapping topics (already, author, call, system).
Shared entity: Judge / Same source domain / Earlier coverage
Both cover Judge; reported by the same outlet (arxiv.org); earlier Judge coverage from 2026-08-16.