Research
DolphinBench Scores Agent Memory on Task Completion and Forces Every Submission to Report Cost and Latency
DolphinBench (arXiv 2609.24971, 21 Sep 2026) argues conversational QA memory benchmarks leak the answer by signaling that a fact must be retrieved and often which one. It instead evaluates memory through agent task completion across three knowledge-work personas with roughly 500k tokens of user messages each, verifying all 200 tasks per persona by running an agent with and without the relevant history and keeping only tasks that succeed with it and fail without. It also requires every submission to report total cost and latency alongside accuracy, closing the loophole where memory systems buy score with unreasonable cost/time tradeoffs.
Source
↳ Follow the thread