Fetching from the wire…
Public story · 2026-08-26 · high
Nine attack variants averaged a 69.5% success rate, and the paper's own runtime defense only cut that to 42.7%.
Why now: As of August 26, agent builders keep adding third-party MCP servers with no standard way to audit what one does after the install review passes.
A Model Context Protocol server can run clean through a conditioning phase, then switch to an adversarial payload once it crosses an interaction threshold, per the paper. Static analysis run before deployment only ever sees the honest phase. A review that clears a server at install time says nothing about what happens after it crosses that threshold.
The authors argue the attack sits outside the threat models people already defend against. It isn't indirect prompt injection, because nothing gets smuggled into content the model reads. It isn't a man-in-the-middle attack either, because the server itself is the adversary, not something intercepting its traffic.
Nine attack variants the authors tested succeeded 69.5% of the time on average, across both frontier and open-weight models. Their runtime defense baselines server payloads during clean trust windows, and even with it running, 42.7% of attacks still got through. Four in ten still land.
That gap matters for anyone wiring third-party MCP servers into an agent. A one-time install review doesn't cover a server that changes behavior later, and the paper's own defense against that still leaves an attacker with close to even odds. The paper doesn't say whether a longer conditioning phase beats the runtime baseline or just outlasts it.
Each link below shares sources, entities, or timing with this story.
OpenAI supports MCP / Shared entity: MCP / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (OpenAI supports MCP); both cover MCP; reported by the same outlet (arxiv.org).
Anthropic released MCP / Shared entity: MCP / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Anthropic released MCP); both cover MCP; reported by the same outlet (arxiv.org).
Linked by a graph relationship (Anthropic released MCP); both cover MCP; reported by the same outlet (arxiv.org).
Claude Code uses MCP / Shared entity: MCP / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Claude Code uses MCP); both cover MCP; reported by the same outlet (arxiv.org).
ChatGPT supports MCP / Shared entity: MCP / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (ChatGPT supports MCP); both cover MCP; overlapping topics (author, server).
Claude Code uses MCP / Shared entity: MCP / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude Code uses MCP); both cover MCP; overlapping topics (attack, server).
Anthropic released MCP / Shared entity: MCP / Same source domain / Earlier coverage / Tension
Linked by a graph relationship (Anthropic released MCP); both cover MCP; reported by the same outlet (arxiv.org).
Claude Code uses MCP / Shared entity: MCP / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude Code uses MCP); both cover MCP; overlapping topics (attack, server).