Terminal-Bench 3.0 Lands With Frontier Models Under 40% and Claude Opus 5 Leading at 43.5%
Terminal-Bench (corroborated by Scale Labs and Turing blog; r/singularity 82up/17c)·high signal
Terminal-Bench 3.0 went live as a harder successor to 2.1, spanning roughly 16 categories from software engineering, security and scientific computing to kernel work, games and debugging, at easy/medium/hard levels, each task programmatically verified. Claude Opus 5 leads the public snapshot at 43.5%, followed by GPT-5.6 Sol at 34.6% and Claude Fable 5 at 34.0%. The value for builders is contamination: the benchmark was assembled through open community contribution under continuous adversarial review and has not yet entered model training sets, so today's numbers are the cleanest agentic-coding signal available.