Fetching from the wire…
Policy2026-09-22 · source-backed
The proposed elements include agent memory and memory access records, actual versus potential autonomy level, and tool usage, none of which existing AI incident frameworks capture. The experts also flagged the reporting pipeline as its own attack surface, since incident records carry leakable context and the infrastructure collecting them becomes a target. Anyone building internal agent incident response should start from this field list rather than adapting a model-incident template. (arXiv 2609.24515)
Each link below shares sources, entities, or timing with this story.
Researchers loaded five systems with a revoked policy and its replacement, then measured retrieval and downstream action across nine policy scenarios, nine models and six defense conditions. Wherever the revocation label was visible to the retrieval layer, the revoked fact cam...
13 public sources consolidated into 9,740 skills (7,505 malicious, 2,235 benign) across 11 harmonized attack categories. Learned text detectors score 0.882-0.932 Macro-F1 under random splits but collapse to 0.653-0.665 source-disjoint. arXiv Three off-the-shelf skill scanners...
A June 26 paper (arXiv:2606.26294) describes a self-improving architecture where the agent and the evaluator that scores it evolve together, specifically to avoid the stagnation of optimizing against a fixed, gameable reward. (arXiv) Anyone building a self-improving harness ha...
Agents get only a CWE description and terminal access, and have to find the implementing files across 500 real vulnerabilities from 290 repositories, six package ecosystems and 147 CWE categories (arXiv 2609.15939). Across 27 language models and four static-analysis tools on a...
Wu et al. name history reliability as a distinct failure mode: trace entries that stay structurally valid and semantically plausible after they stop being authoritative. On Qwen3-1.7B, polluted history flipped 32.1% of decisions correct under the original trajectory, usually v...
arXiv 2607.23444 defeats per-user memory isolation without violating it. Agents routinely embed LTM-retrieved data in tool-call parameters, so a malicious tool exfiltrates memory while every user-ID binding stays intact. SPORE decouples the adversarial command from retrieval a...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.