Fetching from the wire…
Agents2026-09-15 · source-backed
SkillSeam treats the collection, not the file, as the unit of failure (arXiv 2609.13321). Flattening the persistence hierarchy raises loaded-skill tokens 60%. A dangling anchor raises total tokens 64% and costs 3.1pp accuracy. With skill count and context size held fixed, swapping an unrelated control for a synonymous alias drives noncanonical routes from 0/32 to 15/32 and flips half of matched paraphrase pairs. Overlapping lanes raise ownership conflicts from 0/16 to 14/16, bland triggers drive routing conflicts from 3/32 to 30/32 and inflate loaded-skill tokens 3.7x, and one granularity mis-mix costs 12.5pp accuracy. Byte-differenced variants and a one-screen checklist are released.
Each link below shares sources, entities, or timing with this story.
arXiv 2607.08395 observes that persistent agents (long-lived memory, reusable skills, tool-mediated state) have a far larger semantic attack surface than chat assistants, because unsafe content propagates through stored state instead of dying with the session. Nearly all secur...
agentic-kv-cache simulates cross-request prefix caching against real traces, not synthetic ones: 68,266 requests across 393 Claude Code sessions at 64-token blocks, plus Mooncake traces at 512-token blocks. It models prefix-contiguous hits, radix eviction constraints and pinne...
A client receiving isError:true knows something broke but has no machine-readable basis for choosing between fixing an argument, authenticating, waiting, switching tools, or stopping. Auditing 21 safely induced failures across ten reachable MCP servers, typed fields exposed fa...
Skill self-evolution methods revise skill text from execution feedback, but each oracle evaluation needs a full agent rollout, which confines search to patching whatever just failed (arXiv 2609.15396). SkillLift treats ranking as a smoother supervision target than absolute sco...
Production skills are directory bundles where only the root loads at activation and references, schemas, scripts and nested subskills load on demand, so compressing the root misses most of the cost while flattening destroys the progressive-loading boundaries (arXiv 2608.30785)...
arXiv 2609.09218 names two ways a score fails to measure the model: execution-critical decisions get made by a fixed scaffold instead of the model, and the scorer grades output shape instead of task correctness. The repair protocol moves execution decisions to the model, swaps...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.