Fetching from the wire…
Public story · 2026-08-07 · high
FinEvo-Bench tested 120 real financial cases; rubric feedback beat reference-answer scoring in every scaffold.
Why now: Covered in the August 7 briefing on self-evolving agent research.
Four self-evolving scaffolds beat non-evolving baselines on 120 financial tasks, and in Claude Code, skills alone beat pairing them with memory, per Deng et al.
The stakes: evolution cut compliance issues by up to 0.44 per task, an error rate that compounds fast in regulated finance work.
The study, FinEvo-Bench, ran the four scaffolds against paired non-evolving controls on a shared Qwen3.7-Max backbone across 20 business scenes in six financial domains.
Letta posted the highest evolved score, 91.65, and the fewest compliance issues, 0.09 per task. Codex saw the largest gain from evolution, adding 19.37 points over its non-evolving control.
Rubrics beat reference answers as feedback in every scaffold tested, not just Claude Code. The paper doesn't explain why, leaving open whether graded criteria beats answer-matching outside finance tasks too.
The lesson for builders: don't assume stacking memory onto a skill library is free upside. In Claude Code here, it cost points instead of adding them. Test skill-only against skill-plus-memory on your own tasks before wiring both in by default.
Each link below shares sources, entities, or timing with this story.
Codex competes with Claude Code / Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Codex competes with Claude Code); both cover Claude Code, Codex; reported by the same outlet (arxiv.org).
Claude benchmarked against Codex / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Claude benchmarked against Codex); both cover Claude Code, Codex; reported by the same outlet (arxiv.org).
Codex competes with Claude Code / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Codex competes with Claude Code); both cover Claude Code, Codex; reported by the same outlet (arxiv.org).
Claude benchmarked against Codex / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude benchmarked against Codex); both cover Bench, Claude Code, Codex; overlapping topics (claude, code).
Cursor benchmarked against Codex / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Cursor benchmarked against Codex); both cover Claude Code, Codex; overlapping topics (against, beat, claude, code).
Codex competes with Claude Code / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Codex competes with Claude Code); both cover Claude Code, Codex; overlapping topics (against, claude, code, task).
Linked by a graph relationship (Codex competes with Claude Code); both cover Claude Code, Codex; overlapping topics (against, claude, code, task).
Claude benchmarked against Codex / Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude benchmarked against Codex); both cover Claude Code, Codex; reported by the same outlet (arxiv.org).