Research
VoiceMem Splits Speech-Agent Memory Into Informational and Affective Stores, Retrieving in 134 ms
Xie et al. give duplex speech language models a streaming memory architecture with a parallel informational store and an emotional store plus streaming memory I/O, along with a full pipeline for memory-aware training, long-horizon evaluation, and decoupled deployment with interchangeable backends. Under top-5 retrieval the informational side outperforms Mem0 at top-200 by nearly 30 points, and the affective side with short- and long-horizon attribution and dual-node persona modeling improves the aggregate score by 4.29 points over the previous best across three persona benchmarks. Retrieval completes in 134 ms, inside standard VAD latency, so it adds no perceptible conversational delay.
↳ Follow the thread