arXiv Position Paper: 'Coding Benchmarks Are Misaligned with Agentic Software Engineering'
arXiv·medium signal
A June 16 position paper argues that today's coding benchmarks were designed before AI agents existed and therefore mislead: they conflate multiple system components into single scores, penalize valid alternative solutions, and lack the granular feedback signals needed to iterate on agent systems. Useful skepticism for builders reading the current wave of open-weight SWE-Bench leaderboard claims.