Fetching from the wire…
Agents2026-09-15 · source-backed
A different shape of indirect prompt injection defense: predict the likely tool usage at each step, build a local tool prior, flag the planned action when it deviates from what's locally reasonable (arXiv 2609.14987). Reported attack success rates comparable to state-of-the-art defenses with task utility near the no-attack baseline, and an open-source implementation. The framing holds up even apart from the numbers, since content-level filtering keeps losing to obfuscation while action-level auditing sits closer to what you want to prevent.
Each link below shares sources, entities, or timing with this story.
First defense against MCP Tool Poisoning Attacks using a Decision Dependence Graph that correlates LLM attention with tool invocation decisions. 97%+ detection accuracy with zero token overhead. Key insight: behavior-level defenses are fundamentally ineffective against TPA bec...
It synthesizes attack tool-chains in a sandbox, verifies them, renders the verified chain as one natural-looking prompt, embeds state-transition cues in target tool descriptions, and corrects drift mid-run (arXiv 2608.30441). Against Codex, Claude Code and OpenClaw-style harne...
SecOPD fine-tunes a defense using token-level feedback during on-policy distillation rather than the sequence-level signal prior work used. Against PISmith adaptive injections on Qwen3.6-27B it reports 9.0% attack success where Meta-SecAlign, the previous state of the art, sit...
MAFIA (arXiv 2608.03844) targets the two conditions that describe production and that prior attacks failed against: large benign memory pools and active input auditing. It adds placement strategy (probe memory, allocate injection budget, schedule writes to stay retrieval-compe...
Researchers introduced ShareLock, a tool-poisoning attack against MCP that distributes a malicious instruction across several tool descriptions, defeating the assumption that a reviewer reading one tool will catch it. Per-tool review is now insufficient. The attack surface is...
OpenAI Devs announced on August 26 that WebMCP works in the ChatGPT desktop app's built-in browser and in ChatGPT Sites, so ChatGPT and Codex can call a site's declared tools directly. WebMCP is an experimental web standard adding navigator.modelContext to the browser, letting...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.