Fetching from the wire…
Agents2026-09-14 · source-backed
VRL-Bench evaluated every prominent verbal-memory method from Reflexion onward across three models on MiniWoB and WebShop under finite trial budgets. Each improves over memory-free retry in some settings and reduces success in others, and replay experiments show reflection itself can lower success rates by over-exploiting written-down experience at the cost of exploration. Their VEX-squared scheduler was the only update positive in all six settings.
Each link below shares sources, entities, or timing with this story.
SkillPyramid extends Voyager-style libraries with a hierarchical topology plus a self-evolution loop, raising average reward 38% and cutting execution steps 27.7% across ALFWorld, WebShop, and ScienceWorld. The takeaway: don't append every successful trajectory as a new flat s...
arXiv 2608.05906 keeps a dual-polarity memory of verified corrections and observed dead ends for Text-to-SQL repair: 66.34% to 69.79% on Spider, 47.35% to 48.44% on BIRD. Then the authors say the quiet part: paired analysis supports the Spider gain but is weak on BIRD, MERIT i...
The diagnosis in this paper is better than the fix, and the fix is very good. Recurrent memory agents fail at long context, but not for the reason most people assume. The bottleneck isn't capture. It's retention. Retention falls below 30% at 896K tokens because every consolida...
DoCtOR runs automated failure attribution to find the decisive error step and agent, synthesizes what that step should have been via counterfactual reasoning, then asks only that one agent to reflect. Gains over initial success rate: 22% on HotPotQA, 26% on ChartQAPro, 27% on...
IssueLoc-Bench evaluated five explorer models under an identical read-only interface on 499 SWE-bench Verified tasks plus 500 from 153 other repositories, measuring file-finding separately from patching (arXiv 2608.29675). Lower-cost explorers retained 78-94% of reference Hit@...
MemSyco-Bench points out that memory benchmarks test whether memories are correctly stored, retrieved, and updated, never whether the retrieved memory should have influenced the decision at all. Its five tasks check whether agents can reject memory as factual evidence, respect...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.