Research
Widening RAG Context Is an Architectural Trap; Iterative Allocation Gains 16.7-20.5 Points of Portfolio Recall
The paper exposes a 'diagnostic illusion' where standard relevance proxies fail catastrophically on hard negatives, replacing them with a causal leave-one-out probe that isolates what the generator actually relies on. Deploying that probe in a deconfounded factorial grid, the authors argue monolithic context widening is penalized by relevance decay, whereas allocating the same compute iteratively across multiple sequential generations gains 16.7 to 20.5 absolute points of portfolio recall, scaling to 32B models. They package the result as a closed-loop submodular scheduler with an attribution-steered contrastive decoder to force fresh evidence integration.
↳ Follow the thread