Reddit
A context-compressing desktop wrapper cut tokens per correct SWE-rebench answer by 51.1% and lifted solve rate 7/20 to 10/20
Halv published a 20-pair SWE-rebench benchmark on 2026-09-04 using gpt-5.6-luna at medium reasoning through Codex 0.152.0, identical on both arms. Vanilla Codex solved 7/20 and Halv 10/20, with 30.2% fewer total tokens and 24.1% lower total cost, which becomes 51.1% and 46.9% when normalised per correct answer. The authors publish all 40 run records and explicitly say twenty pairs is an early signal that will not generalise across repositories or models, so treat it as a directional result on context compression plus command-output filtering rather than a proven 51%.
↳ Follow the thread