Berkeley measured 21 model-harness pairs: swapping harness moves success rate under 2% but cost up to 5x, and Claude Code's first call carries 10x Pi's context
Melissa Pan, Ion Stoica and Matei Zaharia (UC Berkeley Sky Lab, with Arena) ran seven models across Claude Code, Codex CLI and Pi on 30 sampled tasks each from SWE-bench Lite and Terminal-Bench 2.0, three attempts per task. Harness choice changed success rate by within +/-2% on SWE-bench Lite and +/-5% on Terminal-Bench 2.0, but Claude Code cost 2.0x Pi and 1.6x Codex on SWE-bench Lite by geometric mean of cost ratios, with Fable 5 at 97.8% success for $1.33 in Claude Code versus 96.7% for $0.67 in Pi at nearly identical turn counts (15.3 vs 15.4). The mechanism they identify is measurable before the agent does anything: across all seven models Claude Code's mean initial context is over 10x Pi's, from longer instructions and larger tool schemas. In nine of twelve model-benchmark comparisons a competitor's harness beat the model's own vendor harness, including GPT-5.6 Sol at 83.3% in Pi versus 78.9% in Codex on Terminal-Bench 2.0 at half the cost.
Source
↳ Follow the thread