Fetching from the wire…
Research2026-09-21 · source-backed
This method compresses redundant attempts into nested shortcut trees, reconciles issues across levels using recovery evidence, and summarizes completed branches while expanding unresolved ones, with no post-hoc outcome labels or expert annotation. Under a source-paired protocol that resets environments and contexts for fresh attempts, it gets the highest strict pass rate among non-privileged feedback methods across four recipient models on Terminal-Bench 2.1, improving 7.12 to 15.64 percentage points while the reruns consume 19.0% to 43.6% fewer tokens.
Each link below shares sources, entities, or timing with this story.
Tsinghua's CompactionRL folds summarization into RL rollout collection so the agent learns what to keep when it compresses, optimizing summary and task under one reward. Under fixed 64k–80k windows it adds +5.5 to +7.0 on SWE-bench Verified and +3 to +6.8 on Terminal-Bench ver...
Instead of cloning full teacher demonstrations that mismatch the contexts a student hits at test time, spend a fixed teacher-labeling budget on short continuation rollouts that branch from the student's own trajectories (arXiv). On HotpotQA, ALFWorld, and Terminal-Bench-Dev, b...
Every coding agent leaderboard number you've seen was produced in conditions your security team would reject on sight. Researchers from Accomplish AI and NYU open-sourced Boundary-Bench on August 5. The setup: 12 frontier agent harnesses, roughly 10,000 runs, 89 Terminal-Bench...
arXiv 2607.26598 targets the failure where an agent recovers from an error within an episode but hits the identical failure in later tasks, because post-episode feedback never revises the persistent harness. Guided by a domain-level Evolution-SOP, it writes episodic memory rec...
Same weights. Same tasks. Same time budget. One change to the harness, and fail-to-pass on SWE-bench Verified went from 28% to 49%. The setup in arXiv 2608.26218 is clean enough to be worth describing precisely. 169 SWE-bench Verified tasks, a 20,480-token context window, a fi...
arXiv 2608.08311 describes an agent that continuously rewrites its own tools, prompts, context assembly, and core implementation through reviewed commits. On Opus 5: 86.74% Terminal-Bench 2.1, 90.69% OSWorld-Verified, 0.2301 normalized reward on CL-Bench. It also documents "Ho...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.