BM25 Beats Dense Retrieval by ~20 Points Once a RAG Corpus Passes 10M Tokens, in a 28-Tier Controlled Scaling Study
arXiv / HuggingFace Daily Papers·high signal
A USTC / Metastone Technology / Beijing Academy of Agriculture and Forestry Sciences team posted a corpus-scaling study on 2026-07-30 using EnterpriseRAG-Bench: 511,959 documents and 601M tokens, cut into 28 strictly nested tiers growing 1.25x per level across a ~450x range, holding relevant documents and distractors fixed while only background corpus grows. At full scale BM25 scores 50.5 against dense retrieval's 29.9, and lexical retrieval 'leads at every larger shared tier' past roughly 10M corpus tokens. The design is what makes this credible — it isolates corpus size as the only variable — and it argues directly against defaulting to embeddings for any enterprise RAG index above a few million tokens.