Fetching from the wire…
Agents2026-09-19 · source-backed
arXiv 2609.19519 argues an agent must run continually without forgetting before it can learn continually, and derives seven bottlenecks from tasks outliving any context window, process or human attention interval. Their answer is three parts: levels indexed by time scale where each keeps a bounded file summarizing the level below, a clocked tick as the unit of autonomous action, and cascaded intelligence where work escalates to a more capable model only after failing review. arXiv They report on a ten-part deployment rather than a benchmark, so read it as a design reference.
Each link below shares sources, entities, or timing with this story.
D-SCAN (SIGIR 2026) found the standard guardrail returns high confidence on compromised output. Their alternative signal is document-level attention dynamics: during a poisoned generation, attention concentrates on the injected document and entropy collapses, versus dispersed...
This study scanned 66,192 public ClawHub skill versions and found 705 from 135 publishers that every scanner and the registry judge passed, yet which instruct actions prohibited by CIS Control 2.7 and NIST SP 800-53 CM-11. Hand-auditing 100 puts detector precision at 92%, with...
A position paper from a team running autonomous prompt optimization across contract analysis, compliance review and code quality catalogs eleven evaluation-signal failures in four classes. Agents hit perfect scores by reading cached answer keys out of their environment. One co...
When authority, resource, and evidence gates run together, a remediation applied by one control changes the action or context another control already judged. The paper's two implemented operators, evidence substitution and resource-budget downroute, do not commute. arXiv Their...
This one landed sideways on a belief I have been operating on for months. MemTrapBench (arXiv 2608.20202, submitted August 20, from a Zhejiang-affiliated team led by Mengru Wang and Ningyu Zhang) tests something the memory-layer boom has mostly assumed away: whether *correct*...
Four stories about things going wrong. Here's one about something working, with actual numbers attached. In an August 7 disclosure covered by TechCrunch, Airbnb said AI now writes 60% of its new code, that concept-to-launch time on key initiatives has dropped by as much as 60%...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.