84.7% of Final-Report Errors in an Open Deep Research System Originate at the Orchestrator, Not the Search Agents
Deep research systems pass content and citations between agents like a telephone game, so poor citation recall is hard to attribute. The authors test each agent invocation locally for faithfulness and verifiability against its own inputs, classifying errors as hallucination, uncited input reliance, uncited output or insufficient citations. Applied to three top-ranked open-source deep research systems, almost every agent makes many mistakes except those summarizing a single document, and in AI-Q specifically 84.7% of final-report errors originate at the orchestrator, roughly 31% of those hallucinations and the rest citation mistakes. Two simple interventions guided by the diagnosis raised citation recall 5% with no quality loss.
↳ Follow the thread