Fetching from the wire…
Public story · 2026-09-17 · high
A rival's harness beat the model's own vendor harness in 9 of 12 pairings, and Claude Code's starting context ran over 10x Pi's before work began.
Why now: The seven-model, three-harness comparison entered research coverage on September 17, putting real cost numbers behind a choice most builders make by default.
Three UC Berkeley researchers tested seven AI models across three coding harnesses on two benchmark suites. Swapping the harness barely moved whether a model succeeded, shifting results by at most 2 points on SWE-bench Lite and 5 on Terminal-Bench 2.0. It moved the bill instead, by as much as double.
Melissa Pan, Ion Stoica and Matei Zaharia, working with Arena, ran the three harnesses on 30 sampled tasks each from SWE-bench Lite and Terminal-Bench 2.0. Each task got three attempts, per their HarnessTax study. Claude Code cost 2.0x Pi and 1.6x Codex on SWE-bench Lite, by geometric mean of cost ratios across all seven models.
The gap shows up model by model too. Fable 5 reached 97.8% success for $1.33 in Claude Code, against 96.7% for $0.67 in Pi, at nearly identical turn counts, 15.3 and 15.4. Same result, nearly twice the price.
The researchers trace this to something measurable before the agent takes a single action. Across all seven models, Claude Code's mean initial context ran over 10x Pi's, driven by longer instructions and larger tool schemas. That overhead applies to every task, win or lose.
The pattern held against vendors' own tooling too. In 9 of 12 model-benchmark pairings, a competitor's harness beat the model's own vendor harness. GPT-5.6 Sol reached 83.3% in Pi against 78.9% in Codex on Terminal-Bench 2.0, at half the cost of the Codex run.
Each link below shares sources, entities, or timing with this story.
UC Berkeley's Sky Lab put seven models through Claude Code, Codex CLI and Pi on 30 sampled tasks each from SWE-bench Lite and Terminal-Bench 2.0, three attempts per task, 21 model-harness pairs total. HarnessTax is the result, from Melissa Pan, Ion Stoica, Matei Zaharia and co...
The technique filters terminal output an agent reads one line of, and resolution rates held steady on 50 SWE-bench Lite tasks.
Claude Code v2.1.269 ships an eval harness for plugins and skills with four free graders (regex, tool_used, tool_order, file_exists) and two paid ones (llm, baseline) that require judge model calls.
The update adds a 1M-token context window at $10/$50 per million tokens, plus a way to force one model onto every subagent.
Evolved harnesses lost to a matched-budget sampling baseline once tested on a benchmark the search never saw.
It splits agent composition from runtime adaptation, and its GitHub repos are still active, not archived research code.
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.