Research
Expert Re-Grading Moves GPT-5.6-Sol From 47.3% to 78.7% on HLE-Physics, and Most of the Gap Was Broken Benchmarks
Physics faculty and graduate researchers audited six widely used benchmarks, including ones feeding the Artificial Analysis Intelligence Index, reviewing problem statements, reference solutions, and model responses to separate genuine model errors from grader errors, wrong reference solutions, and underspecified questions. Most cases initially scored incorrect were benchmarking defects, not physics reasoning failures: after correction GPT-5.6-Sol's mean@4 rises from 47.3% to 78.7% on HLE-Physics and 61.0% to 87.2% on CMT-Benchmark, with corrected pass@4 at 94.4% on 54 retained CritPt challenges. UGPhysics, PRISM-Physics, and PHYBench audited subsets also rise substantially.
↳ Follow the thread