Fetching from the wire…
Top 5 · 2026-09-17 · source-backed
UC Berkeley's Sky Lab put seven models through Claude Code, Codex CLI and Pi on 30 sampled tasks each from SWE-bench Lite and Terminal-Bench 2.0, three attempts per task, 21 model-harness pairs total. HarnessTax is the result, from Melissa Pan, Ion Stoica, Matei Zaharia and colleagues with Arena.
Swapping harness moved success rate within ±2% on SWE-bench Lite and ±5% on Terminal-Bench 2.0. Cost moved a great deal more. Claude Code cost 2.0x Pi and 1.6x Codex on SWE-bench Lite by geometric mean of cost ratios. Fable 5 scored 97.8% for $1.33 per task in Claude Code and 96.7% for $0.67 in Pi, at nearly identical turn counts (15.3 against 15.4). One point of success rate, double the bill.
The mechanism is measurable before the agent takes a single action. Across all seven models, Claude Code's mean initial context is over 10x Pi's, from longer system instructions and larger tool schemas. Every turn re-pays for that prefix. And in nine of twelve model-benchmark comparisons, a competitor's harness beat the model vendor's own: GPT-5.6 Sol scored 83.3% in Pi against 78.9% in Codex on Terminal-Bench 2.0, at half the cost.
Put that next to what Patrick Wendell of Databricks posted on September 16. After piloting with about 200 users, Databricks deployed GPT-6 Astra to roughly 3,500 engineers and found it unambiguously better on long-horizon system design work. Total coding spend rose about 60%. Their response was to carve out a dedicated Astra sub-budget so it gets used selectively. Berkeley says harness choice alone swings cost 2x. Wendell says model choice at fleet scale swings the total 60%. Those multiply.
And Steve Yegge, who spent thousands a month on coding agent subscriptions and argued loudest for token maximalism, discontinued Gas Town and conceded he never built anything with it other than Gas Town itself. When the token maximalist reverses, pay attention.
Here's what I'd do Monday. Run your actual task distribution through two harnesses and log cost per completed task, not per token. If the success rates land within a few points, take the cheap one. And measure your initial context before the first tool call, because that's the number you pay for on every single turn. Berkeley's earlier work said the harness barely changes whether you succeed. This is the first work to price that indifference, and it flips the selection criterion from capability to cost.
Each link below shares sources, entities, or timing with this story.
Terminal-Bench 2.1 results (entries dated June 17) put Codex CLI on GPT-5.5 first at 83.4%, Claude Code on Fable 5 second at 83.1%, and Claude Code on Opus 4.8 at 78.9%. The asterisk matters more than the ranking: Fable 5 and Mythos 5 have been export-suspended since June 12,...
Runta published FrontierHarness on September 2 and it's the most directly useful benchmark I've read this quarter, because it controls the one variable everyone conflates. Nine agent harnesses (Codex, Claude Code, OpenCode, Pi, Oh My Pi, DeepSeek Harness, Kimi Code, Exo Harnes...
Everyone spent yesterday arguing about benchmark numbers. Tencent quietly published data suggesting the numbers belong to your infrastructure, not the model. The WorkBuddy Bench leaderboard reports every model under two different agent harnesses — CodeBuddy Code and Claude Cod...
Everyone benchmarks per task. Accuracy on SWE-bench, pass rate on Terminal-Bench, a leaderboard row per model. Together AI ran the experiment sideways: fix the budget at $100, point both models at DeepSWE, and count how much work came out the other end. GLM-5.3 finished 17 tas...
Every number you use to pick a harness comes from public repositories the models may have trained on. Specific Labs built the version that doesn't: tasks drawn from licensed private company codebases, including a 200K-user event app and a fintech processing over 100K bank stat...
I don't care that Grok 4.5 ranks #4. I care that it resolves a SWE-Bench Pro task with an average of 15,954 output tokens where Opus 4.8 spends 67,020. That's a 4.2x efficiency gap, and it lands straight in my monthly bill. SpaceXAI launched Grok 4.5 on July 8, a roughly 1.5T-...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.