Research
Confidence and Perplexity Make Multi-Agent Debate Worse; DEAR Regulates Who Debates Whom Instead
LLMs in multi-agent debate are highly susceptible to blind conformity, and the paper's sharpest claim is that existing individual evaluation methods based on confidence or perplexity fail to track reasoning correctness and can actively exacerbate the problem. DEAR shifts to the group level, quantifying consensus and divergence as group evidence, then running a Selection RL-Agent to choose reference peers and a Behavior RL-Agent to adjust generation behavior, jointly optimized via multi-agent RL as a sequential decision process. Reported gains come with significantly reduced token consumption, which matters for anyone running debate ensembles at cost.
↳ Follow the thread