Fetching from the wire…
Agents2026-07-18 · source-backed
arXiv 2607.11751 describes payloads split across agents so every per-message, per-tool-call, per-step check passes while the assembled object is the attack. The fix is a composition-level verification stage that inspects the final artifact: the written file, the merged diff, the outbound request. All-green step-level checks tell you nothing about the sum.
Each link below shares sources, entities, or timing with this story.
Models navigate to the correct file for 92%+ of required deletions but cut the exact target line only 52% of the time, and 29% of passing patches wrap dead code in a conditional instead of removing it. Grep the diff for newly added if guards around code the task said to delete...
Novel attack class targeting agent *efficiency* not correctness. Triggers cause excessive reasoning steps, dramatically increasing latency without wrong outputs. Agent appears to work but becomes unusably slow. Extends security concerns to denial-of-service via computational w...
If you have a CLAUDE.md, you're in scope. Today. arXiv 2607.14611 (cs.CR, filed July 16) evaluates prompt injection planted in the persistent memory files that agentic coding systems write and re-read across sessions. The researchers tested both Anthropic's Claude Code and Ope...
Thinkingbox is an MCP-compatible sandbox with isolated sessions, full execution traces, and outcome evaluation against terminal backend state, carrying 507 policy-conditioned workflows across retail, hospitality, auto insurance, neobank internal IT and consulting support (arXi...
The attack needs no instruction, trigger, or retriever optimization, just plainly worded false assertions generated in one pass against a LongMemEval corpus. A four-stage screening pipeline that reaches 0.832 recall on indirect prompt injection rejected none of the poisoned me...
The failure they target is specific and under-discussed: a cached error page or a negative price returns in the *expected schema* and gets consumed as fact, unlike a timeout the agent can see. Outcome Monitors check results against contracts mined from task-disjoint traces or...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.