SciCode Was Broken, Not the Models: 263 Defects Found in 65 Problems, Scores Jump From 45-60% to 84-98%
A domain-expert audit of every one of SciCode's 65 problems — a component of the Artificial Analysis Intelligence Index and a standing evaluation in government and national-lab suites — uncovered 263 defects, of which 192 spread across 91% of main problems wrongly rejected correct, instruction-following solutions via non-reproducible gold answers, over-tight tolerances, or self-contradictory specs. Critically, 78% of the score-suppressing defects required specialized physics or math knowledge to spot, not clerical proofreading. Re-evaluating twelve frontier model snapshots on the corrected SciCode-Verified lifts subproblem accuracy from 45-60% to 84-98% and main-problem accuracy from 9-27% to 69-92%, meaning the widely-cited 2026 scientific-coding plateau around 60% was an artifact of the instrument.
↳ Follow the thread