Fetching from the wire…
Research2026-07-18 · source-backed
AgenticSTS (arXiv 2607.02255) uses Slay the Spire 2 as a bounded-memory testbed, isolating how explicit memory layers change outcomes across hundreds of decisions. Enabling a triggered strategic-skill layer took the win rate from 3/10 to 6/10 versus a no-store baseline, using typed retrieval instead of appending raw transcript. The bounded design keeps prompt size flat regardless of run length, which is what makes the ablations clean. 298 documented trajectories and frozen memory snapshots released.
Each link below shares sources, entities, or timing with this story.
The method stores hardware kernel optimization trajectories, with correctness and performance feedback, in an Experience Graph Memory that preserves decision order, observed outcomes, and abandoned branches, then retrieves under a fixed token budget (arXiv 2608.25570). Under t...
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
arXiv 2608.11632 argues storage retention doesn't identify *authoritative* state: unmediated updates by models, tools, and background workers cause stale overwrites and self-authorizing privilege escalation. Untrusted components propose typed changes against an exact predecess...
SodaMem extracts typed events with source attribution and tracks temporal validity so superseded facts are structurally retired rather than competing at retrieval time. 92.8% on LongMemEval-S at $0.00161 per question, median ~18.3k tokens on deepseek-v4-flash, code released. I...
Self-hosted agents read and write their own memory and config to function, which means an attacker can compromise one entirely through legitimate OS system calls with no exploit involved (arXiv 2607.17986). The paper builds a 23-cell attack matrix across Target, Mechanism, Gra...
Engram, a bi-temporal memory engine, scored 83.6% versus 73.2% for a full-context baseline on the 500-question LongMemEval_S benchmark, a statistically significant +10.4 points, while using ~9.6k tokens instead of 79k (arXiv 2606.09900). Roughly 8x fewer tokens and more accura...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.