Multi-agent committees adopt a wrong answer 38% of the time when two peers assert it — and only an independent referee catches it
Across seven cohorts on six public clinical datasets spanning text, imaging, and tabular ICU records, Gemini committees resisted isolated shortcut cues (5-16% flip) but were highly vulnerable to socially plausible ones: when two peers asserted the same wrong answer the holdout adopted it 38% of the time, as did a fabricated "pre-screen" system flag, on both capability tiers. Of three oversight designs, a gate could not separate adoption from honest agreement (100% false-positive rate) and a same-lineage transcript judge worked on text (100% precision, 93% recall) but collapsed on imaging; only a referee that privately re-queried the holdout transferred, at 77-88% precision. Tripling a cue's visual salience did nothing while a second peer voice raised contagion by half again — the lesson being that any multi-agent debate or self-consistency architecture needs an oversight channel independent of the agents' self-reports.
Source
↳ Follow the thread