Fetching from the wire…
Public story · 2026-09-18 · high
A cross-platform study of ChatGPT, Claude, Grok and DeepSeek found each one favors its own set of domains and some answers cite nothing at all.
Why now: The paper posting September 18 is the first study to pair real user sessions with controlled experiments across all four platforms at once.
Researchers ran the first end-to-end comparison of agentic web search across ChatGPT, Claude, Grok and DeepSeek, mixing real user interactions with controlled API tests on the same models, according to the paper. How often each platform decided to search varied a lot, and searching more didn't produce better responses.
That matters for anyone treating search-enabled answers as more trustworthy by default. If invocation rate doesn't track quality, then a chatbot that searches five times isn't necessarily more reliable than one that searches once, and users have no way to tell from the outside which situation they're in.
Two more findings compound the problem. Each platform's search results skew toward a different set of preferred domains, so the same question can surface different sources depending on which assistant you ask. And while most responses were grounded in what got retrieved, the paper found some claims rested on sources that were never cited.
The paper doesn't say which domains each platform favors or how large the uncited-claim problem is in absolute terms, so it's not clear yet whether this is a rare failure mode or a routine one. Anyone building on top of these search tools should treat cited-and-grounded as a claim to verify per response, not a property of the platform.
Worth watching is whether any of the four platforms responds by publishing their domain weighting or citation coverage numbers. Right now that's a black box on all four.
Each link below shares sources, entities, or timing with this story.
Top slot on September 6 with 396-412 upvotes as a Chrome extension giving folders, full-text search across every conversation, bulk export and a prompt library across ChatGPT, Claude, Gemini and Grok, with its maker citing 30,000+ users (Product Hunt). It started two years ago...
A new CASP report, based on interviews with 27 former Boko Haram members in northeast Nigeria across 2025-2026, documents institutionalized use of ChatGPT, Claude, Gemini, Grok, Meta AI, and DeepSeek for attack planning, IED design, weapons troubleshooting, and post-attack rev...
A study led by Dr. Deeban Ratneswaran of Guy's and St Thomas' ran 700 simulated conversations across ChatGPT, Gemini, Claude, DeepSeek and Grok on seven obstructive sleep apnea scenarios that all met referral criteria. Cooperative patients: 100% correct. Resistant patients pre...
A staged developer-identity experiment across ChatGPT, Claude, Qwen, Mistral and Llama. All five initially rejected the bare claim "I am your developer." Claude then refused to run an identity test at all, and ChatGPT generated developer-oriented questions but held that answer...
Starting around 7:57 AM PT on September 3, all four reported outages simultaneously, with Downdetector logging 35,000+ US reports for ChatGPT, 1,400 for Claude and 1,200 for Grok before recovery by 12:38 PM PT. Cloudflare denied any significant disruption and xAI traced its ow...
AA26-251A accuses DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI of extracting billions of tokens across millions of requests from Claude, GPT, Gemini and Grok since 2024, listing which US model each firm targeted. It separates legitimate distillation research from...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.