Fetching from the wire…
Agents2026-09-22 · source-backed
The self-healing harness frames self-modification as admission control: the agent proposes changes to its own operating instructions in an external workspace, and an external runtime gate decides what persists. Candidate rules get provisional authority during evaluation and cross-episode authority only after measured improvement on the triggering failure with no regression beyond a fixed margin on protected cases. Across 16 matched runs on AppWorld, Terminal-Bench and tau²-Bench, 55% of replay-decided proposals helped one case and hurt another. That ratio is the argument for the gate. (arXiv 2609.24130)
Each link below shares sources, entities, or timing with this story.
A ReAct agent scored 77.4% mean accuracy but succeeded on all 5 runs for only 53.0% of tasks, a 24.4-point consistency gap. IBM's Consistency Analyzer replays each decision step with k=5 controlled resampling to find flip-prone decisions without full task replays or ground tru...
"Remember When It Matters" (arXiv:2607.08716) attacks behavioral state decay, where task-critical instructions get buried or evicted on long runs, using a second memory agent that proactively injects reminders into the action agent's context. It gained +8.3pp pass@1 on Termina...
Tsinghua's CompactionRL folds summarization into RL rollout collection so the agent learns what to keep when it compresses, optimizing summary and task under one reward. Under fixed 64k–80k windows it adds +5.5 to +7.0 on SWE-bench Verified and +3 to +6.8 on Terminal-Bench ver...
Here's the number that should sit in every "agents will replace engineers" thread: 15.2%. That's the best model. The mean across 15 frontier models is 4.3%. Tencent Hunyuan's Long-Horizon-Terminal-Bench put 15 frontier models against 46 long-horizon terminal tasks across nine...
Built with 250-plus industry experts, ALE evaluates agents on 960 expert-authored, deterministically-scored workflows across 55 industries. Agents average 26% across all tiers but 2.6% on the "last-exam" tier. The kicker: Codex with GPT-5.5 scores 82% on Terminal-Bench and 0%...
K-Bench scores unlearning across all six channels a ReAct agent exposes, including chain-of-thought, tool calls and tool observations, counting a leak if the secret appears anywhere. When the secret sits in the prompt or retrieval store, the standard benchmarks report no leaka...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.