Metamorphic Testing Reveals LLM Program Repair Benchmarks Confounded by Memorization
arXiv·medium signal
De Koning et al. apply metamorphic testing to diagnose whether LLM-based automated program repair (APR) tools are genuinely reasoning about bugs or simply recalling memorized fixes from training data. By systematically transforming benchmark bugs while preserving their logical structure, the method reveals that several top-performing APR systems show significant accuracy drops — suggesting data leakage, not generalized repair capability, drives reported results.