Fetching from the wire…
Agents2026-09-22 · source-backed
VibeMemBench separates two claims every memory pitch conflates. On 111 coding targets from 90 SWE-rebench V2 repositories with 3,634 history trajectories, directly injecting verified experience raised resolution on four of five held-out solvers by 1.1 to 4.5 points and cut agent steps on all five. But when four existing memory systems had to construct and retrieve that experience themselves, eleven of twelve solver/system pairings lost to the matched memory-off baseline. The experience is in the history. The retrieval layer is what's missing. (arXiv 2609.23570)
Each link below shares sources, entities, or timing with this story.
arXiv 2608.05144 runs Manager, Planner, Engineer and Reviewer roles over persistent project state with *fixed* model weights, self-evolving through runtime state and control policies rather than training. 76.8% on AARRI-Bench, and mature waves use 21% fewer solve-input tokens...
A July 21 paper pairs two near-identical agents: an Explore Agent that inspects untrusted input but holds no tools, and a Safe Agent that takes privileged actions using its own context plus length-constrained hints from the explorer (arXiv 2607.19595). Borrowing from residual...
arXiv 2608.06811 has the plan phase condition memory retrieval while memory-derived trajectory statistics drive stuck detection and replanning, and grounds verification in issue-reproduction verdicts rather than the agent's self-reported completion. +5.0pp over a harness-match...
Coding agents ace correctness benchmarks and flail at repository-level performance work, because bottlenecks hide behind abstraction layers and the agent stops at the first passing patch. PerfAgent wraps an off-the-shelf agent with a profiler-guided, verifier-in-the-loop workf...
A June 16 position paper argues today's benchmarks predate AI agents: they conflate multiple system components into single scores, penalize valid alternative solutions, and lack the granular feedback needed to iterate on agent systems. Read the current wave of open-weight SWE-...
Equips code agents with structured memory built from historical commits, distilling intent-to-code mappings with self-refinement via verification feedback. Using DeepSeek-V3.2 as backbone, boosts SWE-bench Verified from 68.4% to 77.8% — new SOTA. Co-evolution with project hist...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.