Same Model, Different Answers: Enabling Web Search Cut Benchmark Accuracy by 8 Points and 21% of Prompts Gave Inconsistent Results Across Runs
An audit of LLM benchmark methodology (arXiv 2608.06202, Aug 6) ran 401 stratified prompts from BBQ and SafetyBench through both ChatGPT's chat UI and the OpenAI API, with and without web search, collecting 4,812 responses over three repeated runs each. Chat UI responses were less accurate than API responses on both benchmarks with search disabled; enabling web search cut accuracy by up to 8 percentage points and even reversed the direction of the modality trend on one benchmark; and repeated runs of the same prompt disagreed on up to 21% of prompts. Citation grounding and abstention behavior also diverged between modalities — meaning single-modality, single-run accuracy numbers used to argue deployment readiness are measuring one narrow slice of behavior.
↳ Follow the thread