Fetching from the wire…
Research2026-08-09 · source-backed
An audit of LLM benchmark methodology (arXiv 2608.06202, Aug 6) ran 401 stratified prompts from BBQ and SafetyBench through both ChatGPT's chat UI and the OpenAI API, with and without web search, collecting 4,812 responses over three repeated runs. Chat UI was less accurate than API on both benchmarks with search off. Enabling search cut accuracy by up to 8 points and reversed the direction of the modality trend on one benchmark. Repeated runs of the same prompt disagreed on up to 21%. Citation grounding and abstention diverged between modalities too. Every single-modality single-run accuracy number used to argue deployment readiness is measuring one narrow slice.
Each link below shares sources, entities, or timing with this story.
LLM uses OpenAI / Shared entities / Same source domain / Earlier coverage / Tension
Linked by a graph relationship (LLM uses OpenAI); both cover ChatGPT, LLM; reported by the same outlet (arxiv.org).
LLM uses OpenAI / Shared entities / Earlier coverage / Tension
Linked by a graph relationship (LLM uses OpenAI); both cover ChatGPT, LLM; earlier ChatGPT coverage from 2026-07-24.
Simon Willison released LLM / Shared entity: ChatGPT / Shared topic / Earlier coverage / Downstream implication
Linked by a graph relationship (Simon Willison released LLM); both cover ChatGPT; overlapping topics (chat, chatgpt).
LLM uses OpenAI / Shared entity: ChatGPT / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (LLM uses OpenAI); both cover ChatGPT; overlapping topics (chat, chatgpt).
Codex uses ChatGPT / Shared entity: LLM / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Codex uses ChatGPT); both cover LLM; reported by the same outlet (arxiv.org).
LLM uses OpenAI / Shared entity: ChatGPT / Shared topic / Earlier coverage
Linked by a graph relationship (LLM uses OpenAI); both cover ChatGPT; overlapping topics (between, chat, chatgpt).
Linked by a graph relationship (LLM uses OpenAI); both cover ChatGPT; overlapping topics (chat, chatgpt, prompt).
LLM uses OpenAI / Shared entity: LLM / Shared topic / Earlier coverage
Linked by a graph relationship (LLM uses OpenAI); both cover LLM; overlapping topics (audit, between, prompt).