Fetching from the wire…
Research2026-09-01 · source-backed
DEPBENCH collects 203 real-world upgrade tasks across five package ecosystems and five language communities, each containing code-level breakage the upgrade never announced (arXiv 2608.30300). The best completed agent configuration solves 104 of 203, with wide variation across harnesses, models and ecosystems. This is the maintenance half of software work that SWE-bench mostly skips, and half of it not working is a fair description of my own experience letting an agent take a major version bump unsupervised.
Each link below shares sources, entities, or timing with this story.
Traced across 557 SWE-chat sessions (94,813 events) and 33,097 agentic pull requests from AIDev. Agent-facing artifacts account for 60.5% of documentation interactions versus 10.6% for classical technical docs and 1.3% for API references. Consultation is self-initiated 70.2% o...
CodeGrep measures a 30B OpenHands agent averaging 23 rounds and 631K tokens per resolved SWE-Bench Verified issue, much of it grep, glob and view_file. A 14B retrieval agent trained end-to-end with GRPO raises resolve rate to 27.0% from 25.8% while cutting 15% of rounds and 19...
SWE-Touch mines task-critical regions from repair trajectories, builds plausible Counter-Edits that conflict with task completion, and injects them with contextual user messages when the agent reaches that code. Across nine models on SWE-bench Verified, average resolve rate dr...
Coding agents ace correctness benchmarks and flail at repository-level performance work, because bottlenecks hide behind abstraction layers and the agent stops at the first passing patch. PerfAgent wraps an off-the-shelf agent with a profiler-guided, verifier-in-the-loop workf...
While the capability stories pile up, here's the counterweight. As SWE-bench Verified scores cluster near saturation on July leaderboards, an enhanced analysis (SWE-Bench+, on the AIware 2026 benchmark track) found 60.83% of commonly resolved issues contain solution leakage ri...
A June 16 position paper argues today's benchmarks predate AI agents: they conflate multiple system components into single scores, penalize valid alternative solutions, and lack the granular feedback needed to iterate on agent systems. Read the current wave of open-weight SWE-...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.