Fetching from the wire…
Security2026-09-14 · source-backed
K-Bench scores unlearning across all six channels a ReAct agent exposes, including chain-of-thought, tool calls and tool observations, counting a leak if the secret appears anywhere. When the secret sits in the prompt or retrieval store, the standard benchmarks report no leakage while the deployed agent leaks it on 22-86% of queries, with the secret surviving verbatim in the tool-observation channel. When the secret is in the weights, none of twenty published methods demonstrably removes it, and the top-ranked method changes with the base model.
Each link below shares sources, entities, or timing with this story.
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
It synthesizes attack tool-chains in a sandbox, verifies them, renders the verified chain as one natural-looking prompt, embeds state-transition cues in target tool descriptions, and corrects drift mid-run (arXiv 2608.30441). Against Codex, Claude Code and OpenClaw-style harne...
Treats agent control-flow as routing, not reasoning. Uses parallel health monitors + cost-weighted tool graph with Dijkstra shortest-path. When a tool fails, edges reweight and paths recompute automatically. 9 LLM calls vs 123 for ReAct with same correctness. arXiv 2603.01548
Every skill marketplace runs on one assumption: certify each package, and the ecosystem is safe. CompoSkill breaks that assumption by showing composition risk is a path property, not a node property. The attack works black-box. The attacker knows only a role profile. They down...
MetroLLM-Bench is 955 cases across six real metro systems of 37 to 414 stations, requiring structured tool calls and a machine-renderable terminal state. On the 238-case held-out split the 4B student scores 91.3 on Tier 1 against GPT-5.6's 90.6 and 90.0, at Q4_K_M. The gain ov...
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.