EchoNet Benchmarks 9 Open Models on Whether They Notice a Fake Source During Agentic Search
A builder released EchoNet, a benchmark that gives an agent a factual question and a synthetic web to search, seeded with five distinct misinformation shapes: a single fake page, a fake page ranked first, the same fake claim copied across many pages, a loud fake majority surrounding a real primary source, and a genuine update that postdates the model's training. It scores what the author calls epistemic arbitration, the tradeoff between trusting parametric memory and trusting retrieved text, across DeepSeek V4, Qwen 3.8, Nemotron 3 Ultra and six others. This is the failure mode most RAG evaluations skip entirely, and the framing (stubborn models ignore real updates, trusting models swallow fakes) is directly usable when picking a model for a research agent.
Source
↳ Follow the thread