SWE-Bench Pro Exposes Benchmark Inflation: GPT-5.4 Leads at 57.7% vs 80%+ on Contaminated Verified
Morph LLM·high signal
SWE-Bench Pro (1,865 multi-language, uncontaminated tasks) reveals a dramatic performance gap vs SWE-Bench Verified (500 Python-only, contaminated). GPT-5.4 leads Pro at 57.7%, while top models score 80%+ on Verified. A separate analysis found that changing the evaluation harness moves benchmark scores by 22% independently of model choice, making harness selection a hidden variable that can eclipse model selection decisions for coding workloads.