Markets
FrontierHarness Eval: Same Model, Nine Harnesses, 17.5x Spread in Cost Per Task
Runta published FrontierHarness (Show HN, 2026-09-02, 75 points), benchmarking nine agent harnesses (Codex, Claude Code, OpenCode, Pi, Oh My Pi, DeepSeek Harness, Kimi Code, Exo Harness, Hermes) on software engineering and terminal tasks with an identical cold start and fresh checkpoint restore on every run to kill warm-cache bias. Cost per task ranged from $1.05 to $18.34, about 17.5x, while cost per passing task ranged $0.0615 (OpenCode) to $0.2880 (Claude Code), about 4.7x. Codex led on quality at 66.7% pass for $3.47 per task; Exo Harness was cheapest at $1.05 with 53.3% success. The harness, not the model, is now the dominant cost variable in an agentic coding budget.
↳ Follow the thread