Fetching from the wire…
Public story · 2026-07-31 · high
The same models solved IMO problems anyway, the survey says, undercutting tools that debug AI by reading its thinking.
Why now: Quanta pulled these separate findings on reasoning faithfulness into one argument on July 31.
Researchers swapped a correct AI reasoning trace for an incorrect one and the model's final answer didn't get worse, per a Quanta survey published July 31.
That's a problem for anyone building AI observability tools on the assumption that a model's stated reasoning explains its answer. A related finding quantifies it: 30 to 60% of reasoning steps have minimal effect on the output, per the survey.
The piece draws on Melanie Mitchell, Subbarao Kambhampati, Sébastien Bubeck, Tal Linzen and Pavel Izmailov, pulling together separate findings that all point the same direction.
One of those: strings of dots can substitute for human-readable reasoning chains without hurting performance, a swap that leaves benchmark scores unchanged, per the survey.
None of this made the models worse at their jobs. The same models solved IMO problems and improved solutions to 67 math problems, per the survey, while producing reasoning traces that don't reflect how they actually got there.
Each link below shares sources, entities, or timing with this story.
Shared entity: IMO / Shared topic / Earlier coverage / Tension
Both cover IMO; overlapping topics (agent, case, problem); earlier IMO coverage from 2026-03-23.
Shared entity: July / Shared topic / Earlier coverage / Tension
Both cover July; overlapping topics (agent, case); earlier July coverage from 2026-07-30.
Shared entity: July / Shared topic / Earlier coverage / Downstream implication
Both cover July; overlapping topics (agent, case); earlier July coverage from 2026-07-19.
Shared entity: July / Shared topic / Earlier coverage / Tension
Both cover July; overlapping topics (agent, output); earlier July coverage from 2026-07-19.
Shared entity: July / Shared topic / Earlier coverage
Both cover July; overlapping topics (agent, output, problem); earlier July coverage from 2026-07-30.
Both cover July; overlapping topics (agent, anyone, chain); earlier July coverage from 2026-07-30.
Shared entity: Chain / Shared topic / Earlier coverage
Both cover Chain; overlapping topics (chain of thought, problem, reasoning); earlier Chain coverage from 2026-07-02.
Shared entity: July / Earlier coverage / Tension / Downstream implication
Both cover July; earlier July coverage from 2026-07-22; pushes against this story (against).