"Do Not Trust the Benchmark" catalogs five ways general LLM rankings mislead the people choosing models
This perspective paper argues an overall benchmark score is interpretable only relative to the system tested, the questions included and the evaluation conditions, then works through five connected failures: gaps between the evaluated system and the publicly available one, commercial incentives and dependencies in external evaluation, benchmark saturation plus defective tests and contamination, models exploiting the scoring procedure, and the limited relevance of general scores to a user's actual task. It documents specific cases for each and argues for procedures that disclose the tested configuration and validate both the questions and the successful completions. It lands the same week three labs shipped competing benchmark tables, which is the useful context for reading any of them.
Source
↳ Follow the thread