Fetching from the wire…
Research2026-09-04 · source-backed
Mining 3,553 eligible SWE-chat sessions for requirements arriving after the agent has already implemented something, linked to deletion or replacement of prior agent-authored lines. Arrival is followed by roughly 2x the invalidation of matched non-requirement edits, with no decline over the session and no association with operation type. The controlled experiment found delayed disclosure just relocates implementation to after the reveal, while advance warning produced no detected effect. So "tell the agent up front that requirements may change" is advice with no measured support. arXiv 2609.03028
Each link below shares sources, entities, or timing with this story.
Agent Lightning v1.0 (arXiv 2608.17528) inverts the standard agentic RL architecture, and the inversion is the whole point. Normally the training engine owns the environment loop. It drives the agent, collects trajectories, computes rewards. Which means your training setup and...
SWE-NFI builds 188 tasks from merged Python PRs and operationalizes non-functional improvement as 92 executable rules, cleanly separating "tests still pass" from "the code got better." Best agent: 70.0% functional correctness, 0.0-1.3 on structural improvement against a human...
It treats harness improvement as offline learning: diagnose failure traces, generate structured patches that edit the harness as source code, then select updates by validation over mini-batches of failures (arXiv 2608.23041). Gains: 9.0 on GAIA2, 9.6 on SWE-Bench Pro, 10.0 on...
SWE-Touch mines task-critical regions from repair trajectories, builds plausible Counter-Edits that conflict with task completion, and injects them with contextual user messages when the agent reaches that code. Across nine models on SWE-bench Verified, average resolve rate dr...
ProgramBench dropped a benchmark that should make every "AI will replace developers" hot take age badly. The setup: give an agent a compiled executable and documentation, then ask it to architect and implement a complete codebase that reproduces the original program's behavior...
A case study inside a real company migrating 12 features of varying complexity from a corporate ERP written in Visual Basic 6, using one version of the Claude Code agent, evaluating both effectiveness and efficiency (arXiv 2608.28972). Legacy modernization is the use case most...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.