ResearchCalibrated Speculative Decoding: Frequency-Guided Selection Achieves 2.33x Peak ThroughputarXiv·medium signalXBlueskyLinkedInCopy linkConventional speculative decoding suffers from frequent false rejections when draft models produce semantically correct but lexically divergent tokens. CSD adds Online Correction Memory (aggregates historical rejection patterns as rescue candidates) and Semantic Consistency Gating (probability-ratio verification instead of exact token matching). Achieves 2.33x peak throughput speedup over baselines.SourceSource pagearXiv↳ Follow the threadPolicy dependency / Stack layerConverting GUI Trajectories Into Replayable MCP-Style Calls Instead of Unstructured MemoriesarXiv 2609.16635Policy dependency / ContrastA Skill File Can Carry Distilled Memory, Collapsing Two Sets of Retrieval Machinery Into OnearXiv 2609.16669Stack layer / ContrastEmergence World ran 10 agents per world for 16 days and found no frontier model contained an injected attack — one acted on poisoned memory 46 hours laterarXivPolicy dependency / Stack layerEvery Step Passes Its Guardrail and the Workflow Still Violates Policy, and No Step-Scoped Monitor Can Catch ItarXiv 2609.18820Stack layer / ContrastAn Agent's Own Kernel Tuning Reported 10.6x; Against an Honest Baseline It Was 2.03xarXiv 2609.18123Stack layer / ContrastAdding more open models to a multi-agent system usually makes it worse than its own best single modelarXivPolicy dependency / ContrastKV Cache Tiering Buys 73x More Sessions Per GPU, and the Eviction Policy Barely MattersarXiv 2609.16215Stack layer / ContrastGiving a coding agent a refreshing visual map of the repo graph gains 2.4 points on SWE-bench Verified while cutting tokens 5.8%arXiv