Skills
RAG 2026 Benchmark: Recursive 512-Token Chunking Beats Semantic Chunking by 15 Points in End-to-End Accuracy
Vectara's peer-reviewed February 2026 benchmark of 7 chunking strategies across 50 academic papers found recursive character splitting at 512 tokens with 10-20% overlap scored 69% end-to-end accuracy — beating semantic chunking at 54%. Semantic chunking had high recall (91.9%) but produced fragments averaging just 43 tokens, giving LLMs too little context to answer correctly. Key takeaway: semantic chunking's computational cost is not justified by consistent gains on realistic datasets. Start at 512 tokens, measure, adjust.
Source
↳ Follow the thread