Fetching from the wire…
Research2026-09-22 · source-backed
The failure mode: an agent that rewrites its own prompts, control flow and memory shows big in-distribution gains that vanish out of distribution. RRSI constrains both sides, with a temporally annealed budget capping how many edits one candidate bundles, and a selector carrying a critic that screens benchmark-specific proposals plus a pruner that deletes changes too small, too expensive or no longer useful. Across eight benchmarks it gains up to 14.1 points on the evolved split and 4.7 on five held-out ones, and the resulting harness runs on 30% fewer policy tokens. Code at github.com/google-research/rrsi. (arXiv 2609.24972)
Each link below shares sources, entities, or timing with this story.
Headroom compresses tool outputs, logs, RAG chunks, and files before they ever reach the model. It deploys as a library, a proxy server, or an MCP server, and the benchmarks are blunt: 92% token reduction on code search (17,765 down to 1,408) and SRE debugging (65,694 down to...
arXiv 2607.25886 isolates data-centric research capability by fixing the entire post-training stack so only the agent's data strategy varies. Four frontier agents across six benchmarks. Among searches that continued past the best observed score, 78.26% ended on a lower-scoring...
Engram, a bi-temporal memory engine, scored 83.6% versus 73.2% for a full-context baseline on the 500-question LongMemEval_S benchmark, a statistically significant +10.4 points, while using ~9.6k tokens instead of 79k (arXiv 2606.09900). Roughly 8x fewer tokens and more accura...
Distyl AI's IFScale benchmark tested 20 frontier models and found the best one follows only 68% of instructions at 500 lines — meaning one in three rules you write gets silently dropped. Reasoning models (o3, Gemini 2.5 Pro) hold steady through 100-250 instructions before shar...
Skill-α (arXiv 2608.01678) reframes skill generation as RL over sequential edits, decomposing skill construction into individually evaluable changes. The novel signal is a rollback reward that scores each modification by comparing downstream task execution using the original s...
The benchmark has a Controller model receive a structured summary after each coding round and tell a separate fixed Worker agent what to do, verify, or when to stop, which isolates loop guidance from coding ability. Across Controllers the paired reduction in estimated inferenc...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.