Fetching from the wire…
Agents2026-09-10 · source-backed
arXiv 2609.10263 separates what a persistent agent stores from what it uses, because a superseded fact misleads a current-state answer while remaining necessary for a historical query. A retained archive holds everything; a query-conditioned view governs influence, with same-slot replacement links suppressing superseded values in current-state contexts and intent-aware retrieval making older evidence eligible again. A rate-distortion formulation sizes the view to a budget. Ablations show configurations lacking forgetting or query conditioning have the largest deficits.
Each link below shares sources, entities, or timing with this story.
Retrieval returns flat lists of isolated snippets, so agents pick a semantically similar sibling. RepoNav is a post-retrieval interface presenting compact structural cues and candidate targets, guiding on-demand file-structure browsing, and it improves function-level localizat...
LatentMD separates content correctness from boundary correctness in CommonMark fence handling across 9 LLMs and about 37,600 generations. Ablations attribute failures primarily to same-family symmetric-delimiter collisions, not nesting depth, and the problem generalizes to Pyt...
arXiv 2607.29658 attacks the fact that repair agents treat every issue independently and throw away procedural knowledge. STAIR converts historical trajectories into multi-level trees spanning fine-grained diagnostic actions up through high-level strategies, then tailors plan...
arXiv 2608.04804 sends a 7B searcher into the repo first, sandbox-verifies its reproduction claims and strips false ones, then routes to one of four frontier fixers. On the full 266-task Python slice under the official capped budget it solves 159 vs 158 for the best single mod...
Across seven cohorts on six clinical datasets spanning text, imaging, and tabular ICU records, Gemini committees resisted isolated shortcut cues (5–16% flip) but folded to socially plausible ones (arXiv 2608.03744). A fabricated "pre-screen" system flag worked equally well. Of...
A July 21 paper pairs two near-identical agents: an Explore Agent that inspects untrusted input but holds no tools, and a Safe Agent that takes privileged actions using its own context plus length-constrained hints from the explorer (arXiv 2607.19595). Borrowing from residual...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.