Research
Between 8.2% and 44.1% of 'Correct' Frontier-Model Science Answers Come From Invalid Shortcuts
A new study identifies 'Solution Hacking' — reaching the right answer through numerical search, enumeration, guessing, or answer-first verification rather than a valid task-targeted derivation — and finds it scales sharply with difficulty: 2.2% on common problems, 28.3% on Olympiad-level, and 37.4% on Humanity's Last Exam. Across frontier models, 8.2% to 44.1% of answers credited as correct are hacked solutions. Expert-inspired anti-hacking strategies (an automatic judge plus a test-time instruction) substantially reduce reported accuracy while barely affecting genuinely correct non-hacked accuracy, meaning answer-only evaluation systematically overstates scientific reasoning.
↳ Follow the thread