Fetching from the wire…
Agents2026-08-26 · source-backed
arXiv 2608.24306 tests each agent invocation locally for faithfulness against its own inputs, classifying errors as hallucination, uncited input reliance, uncited output or insufficient citation. Applied to three top-ranked open-source deep research systems, nearly every agent makes many mistakes except those summarizing a single document, and in AI-Q specifically the orchestrator generates 84.7% of final-report errors, roughly 31% hallucinations and the rest citation failures. Two interventions guided by that diagnosis raised citation recall 5% with no quality loss. If you're debugging a research pipeline, instrument the synthesis step before the search agents.
Each link below shares sources, entities, or timing with this story.
Shared entity: Applied / Same source domain / Shared topic / Earlier coverage
Both cover Applied; reported by the same outlet (arxiv.org); overlapping topics (agent, applied).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (against, agent, system); pushes against this story (against).
Same source domain / Shared topic
Reported by the same outlet (arxiv.org); overlapping topics (against, agent, deep, research).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (against, agent); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (against, agent); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (against, agent); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (against, agent); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (against, agent); pushes against this story (against).