Execution Feedback Inflates LLM Test-Generation Gains by up to 14.85 Points, and Plain Resampling Wins After the Audit
arXiv 2608.19626 audits the common practice of treating a single accepted program's outputs as ground truth for LLM-generated tests, using 142 development tasks, 114 locked external tasks, 138 held-out tasks, two code models and three seeds. On external inputs where three accepted implementations agree, generated outputs match the panel only 27.79% and 50.12% of the time, and the single-reference oracle inflates measured evolutionary gain by 9.46 to 14.85 percentage points. After correcting for that, equal-budget independent resampling beats mutation-based evolution by 6.01 to 18.83 points, and a real three-round feedback loop is statistically indistinguishable from a density-matched placebo (+0.13 and -0.50 points external).
↳ Follow the thread
No related signals yet.