Skills
Four memory systems, twelve pairings, eleven failures to beat memory-off on real repository coding tasks
VibeMemBench separates two things every memory-system pitch conflates: whether repository history contains useful experience, and whether a memory system can actually deliver it. On 111 coding targets from 90 SWE-rebench V2 repositories with 3,634 history trajectories, directly injecting verified experience raised task resolution on four of five held-out solvers by 1.1 to 4.5 points and cut agent steps on all five. But when four existing memory systems had to construct and retrieve that same experience themselves, eleven of twelve solver/system pairings failed to beat the matched memory-off baseline. The useful experience is in the history; the retrieval layer is what is missing.
↳ Follow the thread