Skills
One plausible-but-wrong document drops deep-research agent accuracy 66–88 percentage points — 'verification inertia' is the failure mode
DRNOISE (arXiv 2607.17291, submitted 2026-07-19) built 100 tasks where the correct answer is supported by two corroborating indirect record chains, then added a single plausible document carrying a conflicting answer. That one intervention caused 66–88 percentage-point accuracy drops in agents that scored strongly on the clean version; the dominant failure was agents retrieving the truthful records but stopping before reconciling the conflict. Generic 'verify your sources' prompts helped but did not close the gap — if you run research agents against the open web, you need an explicit reconciliation step that forces cross-source agreement before an answer is emitted, not just citation capability.
↳ Follow the thread