Agents
Reconstruction benchmark: frontier models recover research ideas from bibliographies alone only 3-15% of the time, multi-agent tournaments hit 23-42%
Reconstruction gives a model only a paper's pre-publication bibliography and asks it to propose the paper's actual research idea, with an anti-leakage protocol using temporal citation cutoffs, anonymous reference IDs and frozen per-paper bibliographies. Across six scientific domains and 643 evaluated papers, seven frontier models achieve just ~3–15% Match rates as scored by an independent LLM judge. A reference-only multi-agent pipeline combining cross-model review with a Swiss tournament over aligned hypothesis slots — and no web search — raises this to ~23–42%, a ~2.4x lift over the best single model, which is a clean quantification of what orchestration buys over raw model capability.
Source
↳ Follow the thread