A 15,336-question study finds multi-agent systems routinely generate the right answer and then vote for a wrong one
arXiv 2608.25937 (26 Aug) separates multi-agent reasoning into candidate generation, peer communication and terminal selection, and holds two of the three fixed to isolate the third, replaying 81,390 fixed candidate pools drawn from 16,278 questions across five benchmarks after mapping judge reliability on 15,336 questions from MMLU-Pro, GPQA, MedXpertQA and MuSR. The headline result is that a correct answer is often already present in the pool while the system converges on a wrong one, a consensus failure the authors call memetic drift. Judge reliability turns out not to be a fixed model trait but to vary with the task, the generator, and how rare the correct answer is, so combining answer frequency with the judge's verdict, changing only the selection rule, lifts accuracy.
Source
↳ Follow the thread