PURPOSE Poisons RAG by Never Contradicting: Non-Conflicting Injections Beat Prior Attacks by 9.7 ASR Points
Post-retrieval conflict resolution is the safeguard that arbitrates among contradictory retrieved passages, and every existing black-box poisoning method trips it by asserting the target answer in frontal contradiction to what the resolver treats as settled. PURPOSE instead extracts query-related facts approximating the resolver's likely reference and grounds a pivot event in them, framing the injection as a consistent update rather than a counter-claim. Across three QA benchmarks, five generators, and three conflict-resolution methods it attains the highest attack success rate in 35 of 45 settings and exceeds the strongest prior attack by a mean of 9.7 ASR points, showing that conflict-detection defenses key on contradiction signals an attacker can simply decline to emit.
↳ Follow the thread