Sources
LMEB: First Long-Horizon Agent Memory Benchmark — MTEB Retrieval Performance Has -0.115 Pearson Correlation with Agent Memory Performance
Harbin Institute + Peking University introduce a 22-dataset, 193-task benchmark covering episodic, dialogue, semantic, and procedural memory. Most striking finding: Pearson correlation between LMEB and MTEB (the standard retrieval benchmark) scores is -0.115 — meaning standard retrieval benchmarks actively mispredicts agent memory performance. Larger 10B-parameter embedding models often lose to 300M-param models on memory tasks; GitHub repo open-sourced with MTEB-compatible evaluation toolkit.
↳ Follow the thread