Two Agent-Memory Projects Now Publish Directly Conflicting Benchmark Claims — MemPalace 96.6% vs agentmemory 95.2%, on Different Splits
MemPalace (57,821 stars, ~512/day, v3.6.0) reports 96.6% raw recall@5 on LongMemEval with no LLM required, 98.4% with hybrid v4 on a held-out 450 questions, LoCoMo R@10 rising from 60.3% to 88.9%, ConvoMem 92.9% and MemBench 80.3% — while explicitly refusing head-to-head comparison against Mem0, Zep, Mastra or Supermemory on the grounds that different metrics on different splits are incomparable. Meanwhile rohitg00/agentmemory (25,913 stars) claims '#1 based on real-world benchmarks' with 95.2% recall@5 and 88.2% MRR on LongMemEval-S, and does publish a rival table citing mem0 at 68.5% and Letta/MemGPT at 83.2% on LoCoMo. Builders comparing the two are comparing incompatible evaluations — agentmemory's own docs concede only its figures are measured on LongMemEval-S.
Source
↳ Follow the thread