Terminal-Bench-Science Launches With 70 Researcher-Authored Tasks and No Model Above 30%
Terminal-Bench-Science / Hacker News·high signal
The v0.1.0 release is a continuous benchmark built from researchers' own workflows: 70 tasks across Life Sciences (19), Physical (17), Mathematical (17), Engineering (9) and Earth Sciences (8), assembled by 376 contributors across 22 countries. Claude Opus 5 leads at 30.0%, followed by GPT-5.6 Sol at 22.4%, Claude Fable 5 at 21.4%, Claude Opus 4.8 at 10.5%, GPT-5.6 Terra at 8.6%, GLM 5.3 at 8.1%, Kimi K3 and Grok 4.6 at 7.1%, and GPT-5.6 Luna at 3.3%. The spread between Opus 5 and Opus 4.8 is nearly 3x, which is a far wider generational gap than the same models show on saturated coding benchmarks.