ResearchSpeculative Speculative Decoding Saguaro 2x Faster Tri DaoarXiv·high signalXBlueskyLinkedInCopy linkParallelizes speculation and verification. 2x faster than speculative decoding, 5x faster than autoregressive. Tri Dao co-author.SourceSource pagearXiv↳ Follow the threadStack layer / Update threadUno Bolts Diffusion Adapters Onto an Autoregressive LLM for 3x Lossless Speedup, and the Weights Already ShippedarXiv 2609.04010Stack layer / ContrastSpeculative Macro Commit Pre-Executes Multi-Action Chains and Cuts Agent Wall Time 44.9% on AppWorldarXiv 2609.03236Stack layer / Update threadSCX Router Scores Model Suitability With No Autoregressive Generation and a Persistent Text-Only KV CachearXiv 2609.02292Stack layer / ContrastIn AI-to-AI messaging, whatever arrives first captures 54% of answers unless verification is enforcedarXivPolicy dependency / Stack layerTyped Provenance Guardrails Block All 19 Unsafe Releases From a Persistent Agent's Autobiographical MemoryarXiv 2609.02127Stack layerVestigeKV Evicts a NoPE KV Cache Using a Vestigial RoPE Branch, Holding 1.00 Retrieval at 32x CompressionarXiv 2609.03949Policy dependencyFP4 Pretraining With E5M3 Block Scales Beats NVIDIA's Transformer Engine Recipe and Runs 21% FasterarXiv 2609.02846Policy dependencySentinel-RL moves the graph reasoning out of the LLM so a SOC agent can operate on a 24M-edge auth grapharXiv