Fetching from the wire…
Security2026-09-21 · source-backed
Micro-Collaborative Poisoning distributes a target claim across multiple individually-plausible documents and evaluates across 108 RAG configurations varying dataset, retriever, retrieval depth, database composition and generator. The effect comes from weak adversarial signals accumulating across sources, so raising top-k and poisoning multiple databases both increase the odds the signals co-occur. Clean database diversity and stronger retrievers dampen it. The practical warning is in the visibility analysis: the attack works with a weaker per-document signature than direct poisoning, so document-level inspection is the wrong detection layer.
Each link below shares sources, entities, or timing with this story.
arXiv 2608.00765 compresses retrieved docs into query-conditioned visual representations, sidestepping the trade-off where hard compression is query-aware but weak and soft compression is strong but needs costly offline encoding. Beats both baselines across varying retrieval d...
Dahal and Xiong target injected documents that are individually benign but create false associations once aggregated, which is structurally invisible to any per-document filter (arXiv 2607.20437). TopoGuard builds a semantic similarity graph over the retrieved set and flags ma...
RAGAS-style evaluation checks correctness against a frozen snapshot, which means routine document updates and corrections can silently break production without moving a dashboard. This ASE 2026 paper defines 11 mutation operators perturbing at both the pre-chunk index level an...
Anthropic commissioned the independent evaluator to test 72 injection scenarios, held out from Anthropic, each run 10 times against Fable 5, Opus 5, and Sonnet 5 as of July 17. Clean sweep. TechCrunch has the details. A third-party held-out eval is a much stronger claim than i...
AgentLSD separates adversarial task contamination from prompt injection: injection needs attacker-supplied instructions, contamination works through non-instructional evidence like fake results and decoy endpoints planted in pages, logs, configs and command output. Six models,...
InceptionRAG fragments the payload into a chain of dormant passages, each benign under isolated inspection, that lead the model to self-deduce the target misinformation through multi-hop reasoning when retrieved together. Across three datasets and three LLMs it exceeds 80% ASR...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.