Fetching from the wire…
Research2026-07-26 · source-backed
KAIST's Yang, Ki, Lee, Choo, and Park target a gap where existing personalization suites either give fully explicit prompts or hand agents pre-abstracted profiles, while real users give underspecified instructions and the context lives in messy history (arXiv 2607.20482). Their PACMem baseline decomposes raw history into factual memories (per-session summaries) and preference memories (recurring behavioral patterns), retrieving from both at inference and consistently outperforming prior memory-based approaches on both task types.
Each link below shares sources, entities, or timing with this story.
Engram, a bi-temporal memory engine, scored 83.6% versus 73.2% for a full-context baseline on the 500-question LongMemEval_S benchmark, a statistically significant +10.4 points, while using ~9.6k tokens instead of 79k (arXiv 2606.09900). Roughly 8x fewer tokens and more accura...
arXiv 2608.05906 keeps a dual-polarity memory of verified corrections and observed dead ends for Text-to-SQL repair: 66.34% to 69.79% on Spider, 47.35% to 48.44% on BIRD. Then the authors say the quiet part: paired analysis supports the Spider gain but is weak on BIRD, MERIT i...
AgenticSTS (arXiv 2607.02255) uses Slay the Spire 2 as a bounded-memory testbed, isolating how explicit memory layers change outcomes across hundreds of decisions. Enabling a triggered strategic-skill layer took the win rate from 3/10 to 6/10 versus a no-store baseline, using...
"Remember When It Matters" (arXiv:2607.08716) attacks behavioral state decay, where task-critical instructions get buried or evicted on long runs, using a second memory agent that proactively injects reminders into the action agent's context. It gained +8.3pp pass@1 on Termina...
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
arXiv 2608.11392 studies what happens when a long-running agent compacts its context: a standing constraint frequently persists as textual residue that no longer governs behavior. Behavioral replay shows models perform the prohibited action far more often with a degraded resid...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.