Fetching from the wire…
Agents2026-09-19 · source-backed
arXiv 2609.20130 diagnoses three failures any agent-memory builder will recognize: episodic memory is badly imbalanced across repositories, more retrieved memory does not monotonically raise success because relevance and redundancy dominate volume, and accumulation is phase-misaligned with piles of reproduction traces and almost no patch or refinement ones. AdaRepair-Mem keeps separate pools for reproduction, localization, patch generation, patch refinement and validation, with coverage-aware fallback to cross-repository memory. arXiv
Each link below shares sources, entities, or timing with this story.
arXiv 2608.05906 keeps a dual-polarity memory of verified corrections and observed dead ends for Text-to-SQL repair: 66.34% to 69.79% on Spider, 47.35% to 48.44% on BIRD. Then the authors say the quiet part: paired analysis supports the Spider gain but is weak on BIRD, MERIT i...
Engram, a bi-temporal memory engine, scored 83.6% versus 73.2% for a full-context baseline on the 500-question LongMemEval_S benchmark, a statistically significant +10.4 points, while using ~9.6k tokens instead of 79k (arXiv 2606.09900). Roughly 8x fewer tokens and more accura...
This benchmark diagnoses the specific step where a medical multimodal LLM goes wrong, instead of only scoring the final answer (arXiv 2606.14697). The stage-wise framing is the transferable idea. Pinpointing the failing reasoning step generalizes to any high-stakes multi-step...
arXiv 2609.20045 audits context compression with paired histories that share the same current answer, receive the same future update, then require different answers. A deterministic frontier selector scored 96/96 strict reveal accuracy on DeepSeek but 82/96 on GLM, a structure...
ERPBench evaluates six screenshot-only computer-use agents against a live reproducible ERP system, scoring against ground-truth database values rather than screen state. Strong general GUI performance does not transfer. The agents reach the right form and save it; the stored r...
Skill self-evolution methods revise skill text from execution feedback, but each oracle evaluation needs a full agent rollout, which confines search to patching whatever just failed (arXiv 2609.15396). SkillLift treats ranking as a smoother supervision target than absolute sco...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.