All 18 Audited Benchmarks Score a Model Name, Not the Route That Actually Served the Request
IBIB (arXiv 2609.10494, 2026-09-09) treats the gap between advertised model identifier and deployed system (weights plus serving route, precision, output contract and harness) as measurement error and gives a protocol that makes it reportable, built on a capability-binding preflight, reliability-inclusive first-pass scoring, and score-blind adjudication. Across eleven systems, two complete runs on identical weights later failed distinct predicates of the finalized binding gate while a third passed, and the advertised identifier exposed neither limit. Serving-arm choice alone moved one declared revision and precision from 77.38 to 82.54, and excluding failed responses from denominators changed the point ordering.
↳ Follow the thread