A 7B Repo Scout Matches the Best Frontier Fixer on SWE-bench Pro at One-Fifth the Cost Per Solve — and the Router Turns Out Not to Matter
SuperScout sends a 7B searcher into the repository first to produce a structured handoff whose reproduction claims are sandbox-verified with false claims stripped, then feeds its hidden states plus task text to a router that picks one of four frontier fixers. On the full Python slice of SWE-bench Pro (266 tasks) under the official capped budget tier it solves 159 versus 158 for the best single model, at roughly one-fifth the cost per solve, with the searcher adding under half a cent of GPU time per task. The honest ablation is the interesting part: always using the cheapest fixer with the handoff ties the routed system, so the verified context handoff — not the routing decision — carries the result, and calibration suggests the handoff redistributes solving ability upward for cheap models while slightly hurting the strongest.
↳ Follow the thread