Pattern: Agent Benchmarks Are Adding a Cost Axis Because Pass Rates Have Converged
Terminal-Bench·medium signal
Terminal-Bench 3.0's leaderboard ranks on resolution rate, cost, and tokens together, and the reason is visible in the numbers it replaces: on TB 2.1, Qwen3.8-Max (86.6), GPT 5.6 Sol (88.8), Opus 4.8 (84.6), and Fable 5 (84.6) sit inside a four-point band. When capability differences shrink to noise, the meaningful question becomes how many tokens a model burns to get there. Expect agent selection to be argued on tokens-per-resolved-task within the next few benchmark cycles, and build your own evals to record token spend alongside pass/fail now.