Automated Benchmark Auditing (ABA): Agentic Framework Systematically Finds Hidden Flaws in AI Benchmarks
arXiv·medium signal
Introduces Auto Benchmark Audit (ABA), an agentic framework that systematically audits individual benchmark tasks to uncover hidden environment dependencies, specification gaps, and brittle evaluation logic that human annotation cannot reliably catch. Applied to SWE-bench, it surfaces issues that inflate or deflate reported scores. Essential context for anyone interpreting agent benchmark results.