Quanta Surveys the Evidence That Chain-of-Thought Is Mostly Theater: Strings of Dots Work as Well as Reasoning, and 30–60% of Steps Have Minimal Causal Impact
Quanta's 2026-07-31 piece pulls together the empirical case against taking reasoning traces at face value, drawing on Melanie Mitchell (Santa Fe), Subbarao Kambhampati (ASU), Sébastien Bubeck (OpenAI), Tal Linzen and Pavel Izmailov (NYU/Google), Weiyan Shi and Pradeep Dasigi (AI2). Three findings anchor it: strings of dots substitute for human-readable chains of thought without loss, 30–60% of reasoning steps have minimal causal impact on the output, and swapping correct traces for incorrect ones does not degrade performance. The uncomfortable framing for anyone building on visible reasoning: the same models solved IMO problems and improved solutions to 67 mathematical problems, so the traces are unfaithful and the answers are still right — which means you cannot debug an agent by reading its thinking.
↳ Follow the thread