Terminal-Bench-Science scores the best agent at 30% across 70 expert-written scientific workflows
Terminal-Bench-Science·high signal
A Stanford-led team from the Terminal-Bench and Harbor project announced Terminal-Bench-Science on 2026-08-28 (115 points on HN), with 70 tasks split across life sciences (19), physical (17), mathematical (17), earth (8) and engineering (9), advised by Ludwig Schmidt and Sanmi Koyejo. Claude Opus 5 leads at 30.0%, GPT-5.6 Sol at 22.4%, Claude Fable 5 at 21.4%, then a steep drop to Claude Opus 4.8 at 10.5%, GLM 5.3 at 8.1% and Kimi K3 at 7.1%. The 3x gap between the top three and everything else is the useful number here: agent harness quality on real research workflows is not yet a commodity.