Expert re-grading finds 60% of CMT-Benchmark questions defective, and GPT-5.6 Sol jumps from 47.3% to 78.7% on HLE-Physics
arXiv 2609.13009, with an accompanying writeup by John Sous dated September 14, had physics faculty and graduate researchers audit six text-only physics benchmarks and found most cases scored as model errors were actually broken answer keys, underspecified problems, or graders rejecting equivalent correct answers. 30 of 50 CMT-Benchmark questions and 21 of 56 CritPt questions contained defects; after correction GPT-5.6 Sol's mean@4 goes 47.3%→78.7% on HLE-Physics and 61.0%→87.2% on CMT-Benchmark, with corrected pass@4 reaching 94.4% on the 54 retained CritPt challenges. For anyone reading eval leaderboards, the practical lesson is that a mid-40s score on a hard science benchmark may be measuring the grader, not the model.
↳ Follow the thread