Fetching from the wire…
Public story · 2026-09-02 · high
A paper describes two attack classes where low-trust input gets folded into high-trust context and stays there after the task that introduced it.
Why now: The paper posted to arXiv in September 2026, naming production harnesses instead of describing the problem in the abstract.
A paper posted to arXiv describes two new attacks on AI agent context windows, tested against 12 harnesses including Claude Code and Codex. The consequences the researchers found range from full agent compromise and remote code execution to denial of service and manipulated tool or skill invocation.
The first, MessageRole Context Privilege Escalation, happens when content from a low-privileged source gets folded into a higher-privileged message role.
The second, Cross-Scope Context Privilege Escalation, happens when that content doesn't clear. It persists beyond the context that first pulled it in, carrying into later turns or tasks, according to the context-privilege escalation paper.
Earlier work titled 'When Context Gets Root' already argued that the harness itself, not the model, is where these systems fail. This analysis is a second independent group reaching that conclusion, and the first to name the shipped harnesses directly.
Unclear from the analysis is whether Anthropic or OpenAI have closed these specific paths already. Equally unclear: whether the 12 harnesses share one root cause or twelve separate implementation bugs. That question decides whether this gets fixed harness by harness or once, at the protocol level.
Each link below shares sources, entities, or timing with this story.
The 34-chapter operations guide says teams conflate instructions, permissions, sandboxing and OS isolation, and that mixup is the top cause of losing control over agent runs.
HarnessRisk ran 128 sandboxed attacks across 14 model/harness setups and found configs that flagged the risk over 90% of the time still let it execute.
Across 46 model endpoints, block rates on the same forged-command test swing up to 47 points between configurations.
Planted skills captured the model's coordinator in 80% of test cases while runtime nearly doubled and task completion stayed unchanged.
A new analysis of AP2 v0.2 found eight high-severity gaps where signed payment mandates don't cover the steps that set up the transaction.
It automates the data-flow, crash-semantics, and commit-history work engineers do by hand.
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.