ResearchSpeculative Speculative Decoding Saguaro 2x Faster Tri DaoarXiv·high signalParallelizes speculation and verification. 2x faster than speculative decoding, 5x faster than autoregressive. Tri Dao co-author.SourceSource pagearXiv↳ Follow the threadStack layer / Threat patternLeakyLMs Fingerprints a Closed Model's Architecture From Token Timing Alone — and Detects Gemini Flash 2.5's ~128K Draft ContextarXiv 2607.20723Stack layerOptimo's Mixture-of-Prompts Hits 3.97x Speedup on Human Code and 13.51x on LLM-Generated CodearXiv 2607.23665Stack layerHiTMS Batches Eight Steganographic Streams Through One LLM Call for 4.3x Throughput and Lower DetectabilityarXiv 2607.23597Stack layerMiniCache Turns Program-of-Thought Programs Into Parameterized Cache Objects for 3.1× Lower LatencyarXiv 2607.20507Policy dependency / Stack layerMolt: NVIDIA's PyTorch-Native Agentic RL Framework Is the Top-Upvoted Paper on HuggingFace This Week at 676 VotesarXiv / HuggingFace Daily Papers (NVIDIA)Threat patternBuild internal agent benchmarks by reverse-engineering tasks from your own commits into colloquial role-played requestsarXiv 2607.20911Update thread / Follow-up threadTRUAV applies distributed multi-agent RL to joint UAV trajectory planning and routing in IoT-enabled vehicular networksarXivStack layer / Threat patternChainWatch scores MCP tool-call sequences with an HMM against a six-stage kill chain, because each call looks benign alonearXiv 2607.19432