Fetching from the wire…
Public story · 2026-09-10 · high
A factorial test on 588 runs shows the model mostly refuses to answer once bad passages take over, rather than inventing facts.
Why now: The paper posted to arXiv on 2026-09-10.
Accuracy for Llama 3.1 8B on a FEVER-derived fact-checking task dropped from 77.9% to 43.5% when researchers poisoned the retrieved passages feeding the model, according to a new arXiv paper that ran 588 factorial test combinations, varying zero to three of three retrieved passages.
That gap matters for anyone shipping retrieval-augmented generation in production. Swap even one of three source documents for something wrong, and the system's answers get a lot less reliable, fast.
The corruption method mattered more than the amount of it. Entity swaps, replacing a name or organization in a passage, flipped the largest share of previously correct answers. Number-based corruption barely moved accuracy while poisoned passages stayed a minority of the retrieved set, then caused a jump in errors once poisoned passages became the majority.
The paper's proxy for unsupported generation, a lexical-overlap check, fell under attack instead of rising. That means the model wasn't inventing content to paper over bad context. It was declining to answer.
The authors used coarse automated labels rather than human review, and they describe the differences between corruption strategies as suggestive, not conclusive. That's a real limit on how far these numbers generalize.
Still, the shape of the failure is useful. If a RAG system degrades under bad retrieval, check whether it's hallucinating or going quiet. Those are different problems with different fixes, one needs better retrieval filtering, the other needs a retry path that surfaces the abstention instead of treating it as a dead end.
Each link below shares sources, entities, or timing with this story.
Hugging Face published its Summer 2026 State of Open Models report on August 14, and one statistic in it went almost entirely unremarked in the coverage. By July 2026, agents rather than humans became the Hub's primary users. Claude Code alone accounted for 44.4% of all agent...
The NYT reported on July 17 that the June 2026 proposal is structured as monthly installments with an early-exit clause for either side, and would sit alongside Anthropic's existing $45B three-year SpaceX GPU deal from May. Meta fell about 6% intraday before closing down 2%. T...
A report on Zuckerberg's internal AI all-hands, including a meeting reportedly interrupted by an employee, surfaced confusion in Meta's direction (Wired). It adds to a run of stories questioning whether the Llama/superintelligence reorg has a coherent plan. For builders depend...
Meta formally pivoted from open-weight Llama to fully proprietary Muse Spark, its first model from the newly formed Meta Superintelligence Labs. No downloadable weights. No self-hosting. Cloud-only private API preview to select partners. More locked down than OpenAI or Anthrop...
Xiaomi released MiMo-V2.5-Pro, a 1.02 trillion parameter mixture-of-experts model (42B active) with 1M token context, fully MIT licensed. In benchmarks, it achieves 63.8% success on agentic tasks using 40-60% fewer tokens than Claude Opus 4.6 or GPT-5.4 for comparable results....
The models are good. The license is the real story. Google released Gemma 4 on April 2 with four variants: E2B, E4B, 26B MoE, and 31B Dense. All built on the Gemini 3 architecture. The 31B Dense variant claimed #3 on Arena AI's text leaderboard, beating models 20x its size. Th...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.