Retrieval-Integration Gap: A Risk Disclosure's Influence on LLM Investment Judgment Falls to Noise at 128K Context While Retrieval Stays Accurate
Holding focal-firm information fixed and varying only unrelated context from 2,000 to 128,000 tokens, the authors show a disclosure's causal influence on an LLM's investment judgment decays to the experimental noise floor even though the model still retrieves it correctly, replicating across model families, judgment tasks, and experiments removing real disclosures from actual 10-K filings. More capable models postpone the gap without eliminating it. Workflow architecture decides the outcome: chunk-and-summarize pipelines evict the relevant information, while a targeted structured restatement placed adjacent to the decision restores influence. The warning for anyone building RAG evaluation is direct, retrieval-based benchmarks will certify systems whose judgments demonstrably ignore what they retrieved.
↳ Follow the thread