Adding more open models to a multi-agent system usually makes it worse than its own best single model
Researchers systematically evaluated 8 model-selection strategies (by size, accuracy and answer diversity) across both before-generation routing and after-generation majority-voting and LLM-as-a-judge architectures on hard scientific benchmarks. Expanding the candidate pool frequently degraded performance below that of the single top-performing base model, leaving a large gap between oracle potential and what the system actually delivers. The one strategy that reliably beat a standalone model was selecting candidates from within a single model family, suggesting heterogeneous multi-model ensembles introduce instability rather than robustness — a direct argument against the reflex of throwing every available open-weight model into an orchestrator.
Source
↳ Follow the thread