MemFuseBench Tests Agent Memory Against Evidence Fragmented Across Apps, Devices and Users
Existing agent memory benchmarks assume a single textual history, which arXiv 2608.18704 (2026-08-19) argues is unrealistic when relevant facts are scattered across applications, devices, users and time. MemFuseBench uses a Scene-to-Sensor pipeline to synthesize controllable scenarios into source-tagged observations, evidence-grounded questions and adversarial distractors, testing temporal reasoning, cross-source fusion and noise robustness. The accompanying MemFuse system keeps source-level evidence in an event-layer atomic memory and groups related events into cluster-layer fused memory inside a causal fusion graph, preserving traceability to original sources; it takes best overall performance under all three LLM settings tested.
↳ Follow the thread