MemArbiter names the "Memory-Action Gap" and gains 20.9 points on ALFWorld under a 500-token memory budget
Posted 2026-08-03 (arXiv 2608.02113), MemArbiter argues the real long-horizon memory problem is not retrieval but post-access failure: information the agent can reach still fails to guide the current decision because it is poorly formed, organized, prioritized, or presented. It decomposes interaction history into atomic items, sorts them into five functional Memory Banks, then combines bank-level demand, item-level relevance, focal/ambient representations, and a temporal presentation gate to control salience per step. On ALFWorld with an open-weight action model it hits 82.8% and 92.5% success under 500- and 750-token per-step budgets, beating the strongest flat-retrieval and flat-recency baselines by 20.9 and 25.4 points, and reduces failed-action repetition and state-action recurrence.
Source
↳ Follow the thread