Fetching from the wire…
Research2026-09-23 · source-backed
arXiv 2609.26457 proposes changes to its own code, benchmarks the modified agents on AI R&D tasks, and keeps only what wins on hidden evaluations. Seven successive improvements in an autonomous run, including a new search policy and context-compressing memory, and the gains generalized to held-out ML engineering, heuristic algorithm engineering and weather forecasting. Gating on hidden evals rather than self-reported scores is what separates this from the self-improvement claims that don't replicate.
Each link below shares sources, entities, or timing with this story.
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
Roesner and Kohno poisoned benchmarks against Darwin Godel Machine, Self-Improving Coding Agent and Hyperagents, using the agent's own self-evaluation loop as the vector. Hyperagents on Sonnet 4.5 self-evolved instructions that disable HTTPS certificate validation on neutral h...
ECP captures agent outputs, tool invocations, and audit context uniformly, with adapters for LangChain, LlamaIndex, CrewAI, and PydanticAI so the same checks run against any of them. arXiv The authors explicitly label it work-in-progress with the method set expected to change....
Danish Foundation Models trained it from scratch on 161 datasets. Across 20 benchmarks spanning English, math and code, and Danish, it beats the original HRM-Text 1B, sets a new Danish state of the art, and competes with Qwen 3.5 4B and Gemma 4 E2B. Weights are on Hugging Face...
SWE-NFI builds 188 tasks from merged Python PRs and operationalizes non-functional improvement as 92 executable rules, cleanly separating "tests still pass" from "the code got better." Best agent: 70.0% functional correctness, 0.0-1.3 on structural improvement against a human...
SpecFirst splits the loop in two: a spec agent probes the binary and fuses observations with documentation into a structured specification, then a separate synthesis agent codes against that fixed reference. Test pass rates rose 6.9-21.3% and binary exploration coverage 9.4-18...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.