InMind Benchmark Exposes the Implicit-Association Blind Spot: Agent Memory Systems Answer 14.4% of Queries the Backbone Gets 84% Right In Context
A four-author paper (arXiv 2607.24368, July 27) names a failure mode every memory-augmented agent has and nobody measures: retrieval assumes the memory you need will textually resemble the query that needs it, which breaks whenever world knowledge is the bridge, as with a stored tree-nut allergy and a macaron request that share no surface cue. On their 125-task expert-verified InMind benchmark spanning ten life domains, with the decisive memory placed directly in context the backbone answers 84.0% of indirect queries, but when the same memory must be retrieved, six vector, graph and agentic memory systems top out at 14.4%, despite recalling those same facts on direct demand at up to 100%. Raising embedding dimensionality eightfold improves target recall for every system and leaves the gap essentially intact, locating the failure in the query-conditioned retrieval interface itself rather than in storage or model knowledge.
↳ Follow the thread