Agents
AttriMem uses token-level attribution to fix the credit-assignment bottleneck in RL-trained agent memory
Submitted July 23, AttriMem augments global outcome rewards with local rewards derived from token-level contributions to the final answer, addressing the problem that outcome-level rewards signal task success but cannot identify which specific memory contents earned it. On long-horizon dialogue QA it reportedly beats retrieval-based, heuristic, and RL-based baselines, generalizes across benchmarks and answer models, and stabilizes RL optimization. No numeric deltas appear in the abstract, so the size of the win is unverified — but the framing that memory-writing decisions need their own reward signal is the useful takeaway.
Source
↳ Follow the thread