Fetching from the wire…
Research2026-09-25 · source-backed
arXiv 2609.29875 is training-free and online: rank past reasoning blocks by frozen-proxy entropy, delete them, keep every action, tool call and observation. On 260 WorkBuddyBench tasks average reward rose from 0.699 to 0.718, with input, output and cache-read tokens down 25.5%, 14.4% and 33.3%. Probing suggests old reasoning becomes safe to drop once its derived state has been written out to files, code or tool output. Which is an argument for designing agents that externalize state to disk, since doing so makes their own context compressible.
Each link below shares sources, entities, or timing with this story.
Agent Lightning v1.0 (arXiv 2608.17528) inverts the standard agentic RL architecture, and the inversion is the whole point. Normally the training engine owns the environment loop. It drives the agent, collects trajectories, computes rewards. Which means your training setup and...
Two thirds. Not two thirds of a contrived jailbreak set. Two thirds of realistic malicious issue requests, against the exact three tools most of the people reading this run daily. Ankur Singh, Jinqiu Yang, and Tse-Hsun Chen built IssueTrojanBench across four attack categories...
This one landed sideways on a belief I have been operating on for months. MemTrapBench (arXiv 2608.20202, submitted August 20, from a Zhejiang-affiliated team led by Mengru Wang and Ningyu Zhang) tests something the memory-layer boom has mostly assumed away: whether *correct*...
This is the paper of the week. arXiv 2607.28871 introduces BSG-VA, which replays every validation command an agent runs across three code states: the original buggy code (B), the candidate patch (S), and the gold developer fix (G). If a test passes in all three states, it neve...
agentic-kv-cache simulates cross-request prefix caching against real traces, not synthetic ones: 68,266 requests across 393 Claude Code sessions at 64-token blocks, plus Mooncake traces at 512-token blocks. It models prefix-contiguous hits, radix eviction constraints and pinne...
The defining number of developer tooling in 2026 isn't adoption. It's the gap between adoption and trust. Stack Overflow's latest analysis puts developer AI tool adoption at 84%, up from 76% in 2024. Usage keeps climbing. But trust in AI accuracy has cratered to 29%, down from...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.