Fetching from the wire…
Skills2026-08-10 · source-backed
PMCoder resolved 25 more SWE-bench Verified cases (+5.0pp) partly by treating issue-reproduction verdicts as the completion signal rather than self-reported success, and by using memory-derived trajectory statistics to detect being stuck and trigger replanning. "The agent said it's done" is the single most expensive assumption in an autonomous loop.
Each link below shares sources, entities, or timing with this story.
arXiv 2608.06811 has the plan phase condition memory retrieval while memory-derived trajectory statistics drive stuck detection and replanning, and grounds verification in issue-reproduction verdicts rather than the agent's self-reported completion. +5.0pp over a harness-match...
This is the most directly usable paper of the day and it does something rare: it bolts onto an existing agent without touching it. Ledger is a deterministic runtime wrapper that distills an agent's completed interactions into explicit state. What has been observed, what has be...
SWE-Touch mines task-critical regions from repair trajectories, builds plausible Counter-Edits that conflict with task completion, and injects them with contextual user messages when the agent reaches that code. Across nine models on SWE-bench Verified, average resolve rate dr...
The most useful AI-productivity dataset I've seen came from a company with every incentive to measure it honestly, because they're 3,500 people trying to run on their own product. The Pragmatic Engineer's July 29 deep dive inside Anthropic reports code output per engineer up 2...
SWEADV built 750 adversarial issue descriptions from 150 SWE-bench Verified tasks, five per task across command execution, deserialization, path traversal, DoS and weak hashing (arXiv 2609.15963). Across mini_swe agents on GPT-5-Mini, MiniMax-M2.5 and DeepSeek-R, adversarial i...
RealSWE builds a six-category information taxonomy and four style dimensions, then compares real prompts from SWE-chat against SWE-bench Verified and Pro. Also: 87% of real prompts are casually written, against 94% of benchmark problems written formally. They release 381 multi...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.