VeriScale: Adversarial Test Suites Expose Overestimation of LLM Code Generation Capabilities
arXiv·medium signal
VeriScale generates adversarial test suites that reveal existing benchmarks systematically overestimate LLM code generation quality by using insufficient positive and negative test cases. The framework scales both the quantity and adversarial quality of tests to evaluate not just functional correctness but formal verifiability of generated code. On existing benchmarks, models that appear high-performing under standard tests show significant degradation under VeriScale's adversarial regime.