Research
More Web Search Calls Did Not Make ChatGPT, Claude, Grok or DeepSeek Answer Better
arXiv 2609.19244 (16 Sep 2026) is the first end-to-end study of agentic web search across four conversational platforms, pairing real user interactions with controlled API experiments using the same platforms' models. Search invocation rates varied substantially across platforms and models, and more frequent searching did not yield better responses. The authors also found each platform's search engine returns results skewed toward its own preferred domains, and that while responses were largely grounded in retrieved results, some claims rested on uncited results, which is an attribution problem for anyone quoting an assistant's sourced answer.
↳ Follow the thread