RAG-NAROK Poisons RAG by Reading What Was Retrieved, Then Writing Documents That Name and Discredit Those Sources
arXiv·low signal
Unlike static corpus poisoning, RAG-NAROK first extracts the identities of legitimate retrieved sources from a transparent RAG pipeline. It then generates 'anchor-specific refutation' documents that name and devalue those sources, exploiting recency and authority biases to steer answers toward a target. Pipelines that expose source citations to users are handing attackers the information they need to target a refutation.