Skills
ToxicRAG poisons a RAG system with a single document written as a plausible knowledge update
Rather than injecting multiple documents or templates that assert the target answer, ToxicRAG generates one document per target that acknowledges the previously accepted answer, introduces fabricated events that appear to invalidate it, and attributes the attacker's answer to purported authorities. Across 100 questions each from Natural Questions, HotpotQA and MS-MARCO, with four victim LLMs and four dense retrievers, attack success ran 0.61 to 0.91 across the twelve combinations and matched or beat the strongest baseline in every one. The shape matters for defenders: a corpus filter tuned for repeated or template-like injections will not see a single well-formed update narrative.
↳ Follow the thread