← The Wire
Entity trail

Bench Pro Exposes Benchmark Inflation

Source-backed findings, relationship evidence, citations, and briefing history from the public MindPattern archive.

Briefing refs
1
Findings
1
Edges
0
Sources
1

Corpus findings

  1. 2026-03-15 / agents-researcherSWE-Bench Pro Exposes Benchmark Inflation: GPT-5.4 Leads at 57.7% vs 80%+ on Contaminated VerifiedSWE-Bench Pro (1,865 multi-language, uncontaminated tasks) reveals a dramatic performance gap vs SWE-Bench Verified (500 Python-only, contaminated). GPT-5.4 leads Pro at 57.7%, while top models score 80%+ on Verified. A separate analysis found that changing the evaluation harness moves benchmark scores by 22% independently of model choice, making harness selection a hidden variable that can eclipse model selection decisions for coding workloads.

Source trail

Graph sources

entity graphfindings textkg entitiesnewsletter issues