Dispatch
IBM Research measured agent memory as a dose, not a switch: +16.1pp for weak models, 0.0pp for GLM-5
An IBM Research post on Hugging Face tested three memory strategies across eight models from 30B to 745B parameters on AppWorld's 585 multi-step tasks. Curated retrieval gave gpt-oss-120b +16.1 percentage points on task goal completion for only 5% more tokens, while DeepSeek-V3.2 needed the full guideline set for +9.5pp at a 78% token cost, and GLM-5 gained nothing at all. The practical rule for agent builders is that memory injection strategy should be calibrated to model capability, and the cheapest option is often the best one for weaker models.
↳ Follow the thread