Fetching from the wire…
Agents2026-09-12 · source-backed
Raw security telemetry arrives faster than a model can read it, single events are ambiguous, and unconstrained model actions carry operational risk. BlueSTAR compresses high-volume telemetry into compact indicators of compromise, reasons over those, and introduces a resilience metric scoring attacker reach, impact on mission-critical assets, and the disruption the defense itself causes. Evaluated on two live enterprise IT/OT ranges against seven attack chains from real intrusion techniques, retaining deterministic playbooks' containment speed (arXiv 2609.11852). Counting your own blast radius as a cost in the objective is the design choice I'd copy.
Each link below shares sources, entities, or timing with this story.
First defense modeling multi-turn indirect prompt injection as temporal causal takeover. Uses counterfactual re-executions at tool-return boundaries to detect when tool outputs steer agent behavior. Evaluated on AgentDojo across four task suites. Builder-ready pattern for tool...
Defense-as-Skill argues pre-install vetting is structurally insufficient because a malicious skill only triggers once a concrete task and workspace state make the unsafe action look useful. SkillSonar runs as an editable skill alongside untrusted skills, checking sensitive act...
MMPIBench pushed a fixed attack set through six visual carriers across 720 runs on six frameworks, five models and four attacker objectives. Visual attacks were attempted in 12.8% of runs but completed in about 1%, with nearly the whole gap closing at the planning step. Extend...
arXiv 2607.26598 targets the failure where an agent recovers from an error within an episode but hits the identical failure in later tasks, because post-episode feedback never revises the persistent harness. Guided by a domain-level Evolution-SOP, it writes episodic memory rec...
The Agent Payments Protocol signs the finished transaction but not the decision behind it, so text in a product description can steer the agent (arXiv 2609.11757). Against the Gemini Flash-Lite models that AP2's sample agents use by default, the attacks fetched another user's...
arXiv 2608.27141 proves a separation result: against an attack whose evidence is fragmented across iterations, any monitor whose safety state resets each trajectory has a true-positive rate equal to its false-positive rate, no matter how expressive it is. A monitor retaining c...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.