Fetching from the wire…
Agents2026-09-17 · source-backed
AgentLSD separates adversarial task contamination from prompt injection: injection needs attacker-supplied instructions, contamination works through non-instructional evidence like fake results and decoy endpoints planted in pages, logs, configs and command output. Six models, 11 web CTF challenges, deterministic trap generation with delivery verification. Agents captured 41% of flags clean, no model solved every challenge, and traps added roughly 20 turns and 2k reasoning tokens even on successful runs. Damage was heterogeneous: some pairs unaffected, others followed decoys or submitted wrong flags. Clean benchmark scores understate real adversarial surfaces, and the token overhead is a budget line.
Each link below shares sources, entities, or timing with this story.
Planting benign-sounding reasoning in an agent's context steers it to adversarial actions while the trace still reads clean, scaling to DeepSeek-R1 (arXiv 2609.15989). Actors paraphrase the injected plan as their own reasoning with no attribution. Giving the monitor access to...
Here's the setup. You build a benchmark with 100 research tasks. Each answer is supported by two independent chains of corroborating records. Clean environment, agents do fine. Then you drop in exactly one plausible-looking document that carries a conflicting answer. Accuracy...
Anthropic commissioned the independent evaluator to test 72 injection scenarios, held out from Anthropic, each run 10 times against Fable 5, Opus 5, and Sonnet 5 as of July 17. Clean sweep. TechCrunch has the details. A third-party held-out eval is a much stronger claim than i...
Data that contradicts the vibe. That's rare enough to lead with. Dipongkor, Baral, Lam and Moran analyzed 4,882 pull requests from five coding agents in the AIDev dataset (532 Java, 4,350 Python), accepted to ICSME 2026. The findings, in order of how much they should change yo...
He warns the payoff only materializes if you structure the codebase so subagents can parallelize, "subroutines but intelligent" (X/swyx). The tactical takeaway: value now comes from architecting your repo for fanout, not just from a better model. Clean module boundaries and is...
Microsoft Threat Intelligence disclosed ChainDrop on August 4: a self-propagating npm worm that poisoned 444 packages across 2,212 versions in under four hours, starting from [redacted] at 150M weekly downloads, plus flat-cache and file-entry-cache. Corroborated by Unit 42, St...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.