One deceptive agent collapses multi-agent truth recovery from 72.5% to 14.2%, and the damage persists after it leaves
The Hi-Agreement framework (arXiv 2608.03421, Aug 4) pairs all-honest collaboration against controlled deception by a single key evidence holder across 120 five-agent object-movement environments where partial observations jointly determine one correct endpoint. Across three LLM-based multi-agent systems, aggregate truth recovery fell from 72.50% to 14.17%. Process tracing showed a single false testimony is adopted more readily than a truthful one, propagates to higher orders, and persists through honest agents even after the deceiver exits the conversation; adding observers without first-hand evidence suppressed incorrect consensus but did not improve truth recovery. Voting and debate do not aggregate evidence robustly — they amplify whoever speaks first with confidence.
Source
↳ Follow the thread