Fetching from the wire…
Public story · 2026-08-24 · high
Researchers permuted the evidence behind 19,520 ReasonBench verdicts and found the judge's stated reasons often didn't hold up.
Why now: The 19,520-case run is large enough that the gap between accuracy and reasoning can't be written off as a small sample.
An AI evaluator scored 98.41% accuracy on ReasonBench, then explained only 54.8% of its own verdicts once the evidence behind them was rearranged. The test covered 19,520 cases. It targets a specific failure: an evaluator can land on the right verdict for the wrong reason, and a standard accuracy score never catches it.
Researchers call the missing piece a judgment receipt, the minimal set of source changes that explains why the evaluator flipped its verdict. It's the difference between the right answer and the right reason. A study testing evaluator judgment receipts found that training on counterfactual examples didn't close that gap.
Static accuracy on an LLM-as-judge setup hides this gap completely. Score a judge only on its verdict, and one that's guessing looks identical to one that reasoned it through. The fix is reporting transformation consistency, how often an evaluator's stated reasons survive when the evidence gets rearranged, next to the accuracy number.
Each link below shares sources, entities, or timing with this story.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage / Tension
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-08-16.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-07-27.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-19.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-18.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-08-21.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-08-17.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-08-12.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-08-07.