Fetching from the wire…
OSS2026-09-04 · source-backed
Terminal agents have produced trajectories at scale while realistic executable environments stay scarce, and environments are what post-training needs, since each can be re-queried into many verifiable tasks while a trajectory is one frozen demonstration. The insight: a trajectory's tool-execution history exposes the structure and contents of the environment it ran in, so replaying recorded file operations restores each file to its pre-modification state, yielding a partial workspace a completion agent fills in. It scales along breadth, mining directional dependency relations between environments for cross-workspace queries, and depth, extending single-turn queries into longer interactions. Top paper on Hugging Face Daily Papers at 115 upvotes. arXiv 2609.04148
Each link below shares sources, entities, or timing with this story.
Here's the number that should sit in every "agents will replace engineers" thread: 15.2%. That's the best model. The mean across 15 frontier models is 4.3%. Tencent Hunyuan's Long-Horizon-Terminal-Bench put 15 frontier models against 46 long-horizon terminal tasks across nine...
Instead of cloning full teacher demonstrations that mismatch the contexts a student hits at test time, spend a fixed teacher-labeling budget on short continuation rollouts that branch from the student's own trajectories (arXiv). On HotpotQA, ALFWorld, and Terminal-Bench-Dev, b...
The top-upvoted paper on Hugging Face Daily Papers this week at 167 upvotes runs a parallel "left brain" for factual memory and "right brain" for affective attribution and dual-node persona modeling, with streaming memory I/O and interchangeable backends. At top-5 retrieval th...
Same weights. Same tasks. Same time budget. One change to the harness, and fail-to-pass on SWE-bench Verified went from 28% to 49%. The setup in arXiv 2608.26218 is clean enough to be worth describing precisely. 169 SWE-bench Verified tasks, a 20,480-token context window, a fi...
arXiv 2608.08311 describes an agent that continuously rewrites its own tools, prompts, context assembly, and core implementation through reviewed commits. On Opus 5: 86.74% Terminal-Bench 2.1, 90.69% OSWorld-Verified, 0.2301 normalized reward on CL-Bench. It also documents "Ho...
Tsinghua's CompactionRL folds summarization into RL rollout collection so the agent learns what to keep when it compresses, optimizing summary and task under one reward. Under fixed 64k–80k windows it adds +5.5 to +7.0 on SWE-bench Verified and +3 to +6.8 on Terminal-Bench ver...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.