Fetching from the wire…
Top 5 · 2026-09-03 · source-backed
Neither report cites the other. Both went up September 2. They attack opposite ends of the same pipe and both find it broken.
Haus Research took the retrieval end. They asked perplexity/sonar and sonar-pro ten question templates about 210 technology companies, 310 questions total, then fetched every cited URL and checked whether the page held the figure it was cited for. Of 1,826 citations attached to numeric claims across 2,915 unique URLs, 34.7% pointed at a page that either wouldn't open or contained none of the sentence's figures. Scored per claim instead of per citation, 14.4% of 872 claims fail. Dead links were only 1.3% of the problem. Paywalls were 16.1%, and readable-but-unsupporting pages another 16.1%. Datasets and scripts published alongside.
Trellner Research took the supply end. Querying the same models across 380 software categories, they found worldmetrics.org, wifitalents.com and gitnux.org together hosting 215,128 templated "best [category] software" pages, drawing 181 of 7,534 collected citations. All three domains registered between December 2023 and May 2024. Shared Cloudflare nameservers. Identical page templates. Their homepages carry the HTML title "Facts & Grounding Page" with meta descriptions advertising machine-readable verified facts. That is a page built to be cited by a model rather than read by a person, and it says so in its own metadata. Trellner also found 59.8% of sources behind grounded AI software recommendations sit outside the top 100,000 sites.
Put those together and the picture is ugly for anyone who assumed answer engines are replacing G2 and Capterra with something more trustworthy. The supply side has 215,128 machine-authored pages farming the channel. The retrieval side gets the number wrong on a third of numeric citations. Haus found the "entry price" question type passed only 62.6%, which means when a buyer asks an answer engine what your product costs, there's better than a one-in-three chance the answer traces to a page that doesn't say that.
If you sell software, the instruction is specific and boring: make your pricing page machine-fetchable and unambiguous. No JavaScript-rendered price. No "starting from" without a number. No pricing that only exists inside a PDF or behind a contact form. If the model can't fetch a clean figure from you, it will cite a templated aggregator page that made one up, and you will not know it happened.
I don't think this kills answer-engine discovery. Search had the same problem for fifteen years and content farms got mostly beaten back. But the mechanism that beat them, human users noticing the page was garbage and bouncing, doesn't exist when the consumer is a model that reads the page once and never comes back. I don't know what replaces that feedback loop, and neither report proposes one.
Each link below shares sources, entities, or timing with this story.
OpenAI Devs announced on August 26 that WebMCP works in the ChatGPT desktop app's built-in browser and in ChatGPT Sites, so ChatGPT and Codex can call a site's declared tools directly. WebMCP is an experimental web standard adding navigator.modelContext to the browser, letting...
guillaumemeyer/watermarks-remover appeared August 11 and hit 634 stars and 61 forks in 24 hours, targeting provenance marks from Anthropic, Google's SynthID-Text, OpenAI, and Kirchenbauer-style implementations across three layers: invisible Unicode hygiene, statistical token-s...
Two thirds. Not two thirds of a contrived jailbreak set. Two thirds of realistic malicious issue requests, against the exact three tools most of the people reading this run daily. Ankur Singh, Jinqiu Yang, and Tse-Hsun Chen built IssueTrojanBench across four attack categories...
Google DeepMind shipped computer use as a built-in tool inside Gemini 3.5 Flash on June 24, collapsing the standalone Oct-2025 computer-use model into the same model you already use for function calling, Search grounding, and Maps. One agent sees a screen, clicks, types, and l...
Google calls it the biggest Search change in over 25 years, rolled out through mid-June. Search answers the query directly and builds a page around the answer rather than returning a list. For anyone shipping content, this is a structural hit to click-through economics. The an...
Check your GitHub Copilot settings right now. As of April 24, GitHub's updated privacy policy flipped the default for all Copilot Free, Pro, and Pro+ users: your interaction data, including prompts, suggestions, and code snippets from your context, now trains AI models unless...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.