Conversational framing exposes agent memory retrieval gaps that QA-style benchmarks hide
LOCOMO-CONV rebuilds LoCoMo into a conversational memory benchmark with four query styles (dialog, implicit, counterfactual, composed) and evaluates five representative memory systems on both retrieval recall and end-to-end response quality, rather than the QA probing existing benchmarks use. Implicit and composed queries expose substantial retrieval gaps, which multi-facet query rewriting narrows for raw-turn memory but not for abstractive memory. Strong retrieval also does not translate into response quality, and implicit queries show 'silent grounding' where memory improves contextual grounding without ever surfacing the gold fact; they release supportive_memory annotations covering conversationally useful context beyond the original gold evidence.
Source
↳ Follow the thread