Fetching from the wire…
Infra2026-09-17 · source-backed
arXiv 2609.17983 formulates stale KV-cache repair as budgeted recomputation and compares training-free position-selection policies on a factual RAG benchmark with matched direct and derived edits. All policies handle direct edits; derived edits separate them, and at the primary budget a contiguous edit-local window recovers at least 0.94 of the post-edit answer margin, substantially beating attention-based, KV-deviation and structural selectors. The mechanism is that scattered positions inherit surrounding staleness even when they look important under clean-state transplantation, and the advantage depends on adjacency, largely disappearing when the answer text sits downstream.
Each link below shares sources, entities, or timing with this story.
arXiv 2608.00765 compresses retrieved docs into query-conditioned visual representations, sidestepping the trade-off where hard compression is query-aware but weak and soft compression is strong but needs costly offline encoding. Beats both baselines across varying retrieval d...
A July 2 evaluation tested semantic chunking against simple approaches on long structured academic theses using RAGAs, and the sophisticated method didn't win. Performance varied more with document formatting, preprocessing, and query type than with chunking strategy. The auth...
arXiv 2608.04756 observes that post-retrieval conflict resolution catches existing black-box poisoning because every prior method asserts its target answer in frontal contradiction to settled context. PURPOSE extracts query-related facts approximating the resolver's likely ref...
RAGAS-style evaluation checks correctness against a frozen snapshot, which means routine document updates and corrections can silently break production without moving a dashboard. This ASE 2026 paper defines 11 mutation operators perturbing at both the pre-chunk index level an...
PCAS: Policy Compiler for Secure Agentic Systems — The first paper to provide measured enforcement results for agent policy compliance (48% to 93%). Uses dependency graphs and Datalog-derived policy language with a reference monitor intercepting all actions. Three case studies...
TraceJudgeBench audits citation artifacts in RAG evaluation across GPT-5.5, Claude Sonnet 4.6 and DeepSeek V4-Flash. Stronger anti-citation prompting cut worse-cited wins from 50.5% to 0%, but some operating points start converting validated moderate-gap decisions into Ties we...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.