Expanding a RAG Index Flipped 10.25% of Answers While Accuracy Moved Only 1.5 Points
A retrieval-augmented QA system can return different answers after an index expansion even with model, prompt, retrieval policy, evidence depth, and generation controls held fixed, and aggregate accuracy hides it when gains and losses cancel. The Snapshot Compatibility Audit estimates excess churn by subtracting same-snapshot repeat disagreement from cross-snapshot disagreement; expanding a frozen FineWeb prefix from one to seven shards produced 6.44 points of normalized-exact and 10.25 points of blinded-semantic excess churn on a preregistered 400-question Natural Questions study, while exact-match accuracy changed by only -1.50 points. A post-hoc pass found repeat-stable semantic flips on 40 of 400 questions.
↳ Follow the thread