Fetching from the wire…
Agents2026-07-29 · source-backed
arXiv 2607.25255 names the real multi-agent failure: a harmful objective gets fragmented into locally plausible subtasks, so no single agent ever sees enough to refuse. SafeFlow attaches structured semantic taints to root requests, propagates them through a dynamic collaboration graph, and validates at the workflow level before any irreversible action. Across four benchmarks it cuts attack success below both undefended baselines and external defenses while keeping benign completion high.
Each link below shares sources, entities, or timing with this story.
Every guardrail I've used adjudicates the current action, which means it can only stop the last step of a plan it never saw coming. JANUS trains a guard on partial trajectories to anticipate safety-relevant futures, then judges from both the observed prefix and the forecast, o...
5 novel attack types (intent hijacking, tool chaining, task injection, objective drifting, memory poisoning) across 28 environments. Key finding: single-turn defenses fail against multi-turn adversarial strategies. (arXiv 2602.16901) ---
arXiv 2608.09902 wraps all 22 boss encounters of Dark Souls: Remastered in a containerized Gymnasium-style benchmark where each step is a real action against the running game. On DSLE-5, an expert system and an evolutionary baseline beat only the tutorial boss (63% and 43% pea...
When an agent consolidates an external observation into long-term memory, attach platform-controlled metadata recording the source's trust level, then gate tool execution by matching action risk against supporting-memory authority. Laundered memories hit a 1.000 attack success...
arXiv 2607.26998 flips the pentest agent's observation-action loop against it, replacing static honeytokens with a trajectory-adaptive policy that constructs new decoy artifacts conditioned on the agent's interaction history, folding validated ones into a factually consistent...
Microsoft's July 23 release targets a genuine gap: harness-based agents like Claude Code and Codex drive multi-turn reasoning, tool use, and external system access but were hard to train end-to-end with standard open RL infrastructure. The trick is decoupling training from inf...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.