Research
Harness-Layer Auto-Research Cut Agent Token Traffic 44.7-49.0% at Equal Task Performance
SoL-Pi (arXiv 2609.20519, 17 Sep 2026, 38 HuggingFace upvotes) scales automated research loops over agent harness designs rather than over models, keeping four mechanisms that survived selection: action execution, context compaction, observation handling and delegated reading. On the 51-task EdgeBench evaluation with GPT-5.6 Sol and Opus 5, it matched Pi's performance while cutting recorded token traffic 44.7-49.0% and API cost by roughly a third. The authors put the saving at $8.75-$13.50 per hour against native Codex and Claude Code harnesses, which makes harness configuration, not model choice, the cheaper lever for anyone running agents continuously.
↳ Follow the thread