Fetching from the wire…
Security2026-09-07 · source-backed
A paper reframing indirect prompt injection as test-time search builds an agentic attacker doing environment reconnaissance, structured strategy reasoning and adaptive evaluation against victim feedback (arXiv 2609.04495). More attacker compute consistently improves both vulnerability discovery and exploitation, and ablations show explicit strategy management is what stops redundant search from eating the larger budget. The consequence for anyone reading red-team reports: a resistance number published without an attacker compute budget tells you nothing, and your measured resistance falls as budgets rise.
Each link below shares sources, entities, or timing with this story.
On a verifiable protein-function characterization task routed across tools, model choice swamped federation topology, RL-versus-LLM harness, and prompt expertise: Opus at roughly 92 to 94%, o4-mini at 40 to 50%. Federation across institutional boundaries cost almost nothing (a...
arXiv 2607.12227 (Wang et al., incl. Hajishirzi, Tsvetkov, Dasigi) finds two methodological holes in the self-improving-agent literature: methods are never compared against simpler baselines at matched compute budgets, and final performance gets reported on the same public ben...
A fleet evaluation across 46 endpoints from six vendors found a recognition-enforcement gap: source-format features are linearly decodable from activations and models verbally identify forged authority when asked, but some configurations still emit the conflicting tool call. A...
Across 30 models from three families, verbalized confidence compared against logits-based confidence on 8 classification tasks and semantic entropy on 2 generation tasks: instance-level association is weak on average, improving only on easier items and stronger base models. In...
arXiv 2607.24174 (July 27) generated adversarial log entries from real attack traces and got multiple state-of-the-art LLMs to classify traces containing clear indicators of compromise as benign. The defensive gift: the natural-language explanations emitted alongside the class...
Add a PromptFoo eval suite that runs every prompt change against fixed test cases across multiple models and fails the build on regression (Lakera). Prompt edits become reviewable diffs instead of blind tweaking. Exactly like unit tests, because that's what they should be.
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.