FinEvo-Bench measures whether agents actually learn from experience — Letta scores 91.65, Codex gains the most at +19.37
Deng et al. (arXiv 2608.06144, submitted August 6) built a longitudinal benchmark of 120 real-case-grounded tasks across 20 business scenes in six financial domains, where each scene shares an institution-provided procedure and a manually reviewed rubric, so later tasks can benefit from earlier experience. Four self-evolving scaffolds on a shared Qwen3.7-Max backbone ran three interleaved task streams against paired non-evolving controls: Letta posted the highest evolved score (91.65) and fewest compliance issues (0.09 per task), while Codex showed the largest self-evolution gain (+19.37); evolution lifted scores 9.33–19.37 points and cut compliance issues 0.12–0.44 per task across the board. Two findings matter for builders: in Claude Code, skill-only evolution beat both memory-only and combined memory-skill evolution, and rubric feedback beat reference-answer feedback in every scaffold.
Source
↳ Follow the thread