Fetching from the wire…
Public story · 2026-07-23 · high
Escaped quotes and curly dollar signs planted in sender-name fields fooled six frontier models, beating purpose-built defenses half the time.
Why now: The findings come from arXiv 2607.05120, posted in July 2026 by Seoul National University, UIUC, and Largosoft.
A new technique hides fake quote marks and dollar signs inside data an AI agent already trusts, hijacking its behavior, per a Seoul National University, UIUC, and Largosoft paper.
That's a problem for anyone shipping agents that read text they don't fully control, like emails or scraped web pages. The technique succeeded 31 to 43 percent of the time on structured data and beat purpose-built prompt-injection defenses up to half the time.
The technique is called probabilistic delimiter injection. Attackers hide escaped quotes, curly quotes, and dollar signs inside fields an agent already trusts, like sender names or button IDs. The model reads those characters as structural syntax, the way it would read real formatting. A strict parser would just see plain text.
On webpage data, success rates ran as high as 100 percent, per the paper. Classic injection attempts, the obvious ones, got caught almost every time. This technique reads as ordinary data instead.
Two defenses actually worked, per the paper. Assigning unguessable random IDs to page elements cut success from about 49 percent to 29 percent, a real reduction but not a fix. Full data-source lineage verification, tracking where each piece of input came from, eliminated every tested attack.
A related benchmark on agent security found single-turn attacks mostly fail, while attacks that adapt across many rounds succeed far more often. Delimiter injection looks like a single-turn version of that same pattern, dressed up as a formatting quirk instead of an attack.
Each link below shares sources, entities, or timing with this story.
An attacker stole an AI agent's signing keys through email injection in under five minutes, per a prior incident this design cites.
A training-free fix called ChannelGuard held steady across three model backends, filter or no filter, blocking every tool-poisoning attempt.
A new analysis of AP2 v0.2 found eight high-severity gaps where signed payment mandates don't cover the steps that set up the transaction.
It automates the data-flow, crash-semantics, and commit-history work engineers do by hand.
A proposed provenance gate cut unauthorized high-risk actions to zero after the attack itself hit a 1.000 success rate in tests.
Within 48 hours, three unrelated sources landed on the same structural problem from three directions.
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.