Early-outcome predictors for SWE-bench agents mostly transfer, but gpt-5-mini and claude-opus-4.6 heads carry persistent calibration gaps
arXiv·low signal
Paper 2609.25647 audits trajectory-based early-stopping predictors across agents using public SWE-bench Verified trajectories. Broad transfer failure was not supported (median corrected gaps 0.018 and 0.039), but two target/head pairs, gpt-5-mini/SUCCESS and claude-opus-4.6/FAILURE, kept gaps of about 0.14. Recalibrate on each target agent before killing eval runs early.