← The Wire
Entity trail

ReasonBench

Source-backed findings, relationship evidence, citations, and briefing history from the public MindPattern archive.

Briefing refs
1
Findings
1
Edges
0
Sources
1

Corpus findings

  1. 2026-08-24 / skill-finderAn evaluator scoring 98.41% recovered only 54.8% of valid judgment receipts when its evidence was permutedJudgment receipts are the minimal set of source changes that explain why an evaluator flipped its verdict, and they separate getting the right answer from getting it for the right reason. On ReasonBench, 19,520 cases with verifiable receipts across policy and logical reasoning, one model with 98.41% frozen-test accuracy recovered only 54.8% of valid receipts once sources were permuted, and counterfactual training alone did not close the gap. If you run an LLM-as-judge in a pipeline, this says report transformation consistency next to accuracy, because static accuracy hides the failure.

Source trail

Graph sources

entity graphfindings textkg entitiesnewsletter issues