Fault Localization Loses to Blind Resampling 3 to 40 in a Placebo-Controlled Code Repair Study
Three arms were applied to the same failed candidate: blind whole-solution resampling, spectrum-based localization followed by suspect-span infilling, and same-length infilling at a random disjoint span. Across three frozen 26-32B models, three benchmarks, and 488 failing candidates, localization was rarely even available (only 9.0% of failing candidates expose a failing public test with a usable spectrum), and where it was, localized infilling lost to blind resampling at matched attempt count 3:40 (p = 3.0e-9), replicating at -11.3 points in a fourth model from a third family. Sixteen localized attempts reach 6.8% while one blind attempt already reaches 10.1%, because infilling reproduces the removed span verbatim 48.9% of the time.
↳ Follow the thread