Global LLM Leaderboards Are Misleading: Top 50 Models Are Statistically Indistinguishable
arXiv·high signal
Analysis of ~89K pairwise comparisons across 52 LLMs in 116 languages from Arena shows that 2/3 of decisive votes cancel out, and pairwise win probabilities among the top 50 models are at most 0.53 — essentially a coin flip. The standard Bradley-Terry ranking creates an illusion of meaningful ordering where none exists. For practitioners: stop chasing leaderboard deltas and benchmark on your actual workload.