Risa Uses MoE Router Traces to Pick Patches, Lifting SWE-bench Verified Resolve Rate From 44.9% to 48.2% With No External Judge
Test-time scaling for software agents is hard because patches have no canonical answer form and sibling actions from a shared prefix are correlated. Risa reads native mixture-of-experts router traces as a behavioral role signal, using them within trajectories to encourage exploration then converge during patch commitment, and across separately sampled trajectories to select a final candidate by agreement at informative patch positions. On SWE-bench Verified it raised the macro-average resolved rate from 44.9% under uniform sampling to 48.2% on the gpt-oss family, matching text consensus without answer-string matching or selection-time test execution, and transferred to Qwen3.6 on the full 500-task benchmark.
↳ Follow the thread