Self-Audit of an LLM Eval Finds Only the Bottom of the Leaderboard Is Stable, and Half the Endpoints Vanished in 10 Weeks
arXiv·low signal
arXiv 2609.30074 bootstraps its own eight-model ranking, covering 8B to 675B models with caching disabled. The two worst models held rank in 99% and 86% of replicates, the middle four in only 27-48%, and the top two in 68%. Two defensible rules for merging runs changed four of eight rows. Four of the eight API endpoints were withdrawn within ten weeks, so the study can no longer be rerun. Small-prompt-set leaderboards should report rank stability.