SourcesMiniCPM-SALA: Hybrid Attention for Efficient Long-Context (8600 HF upvotes)HuggingFace Daily Papers·high signalXBlueskyLinkedInCopy linkHighest-upvoted HuggingFace paper. Hybrid sparse-linear attention for efficient long-context modeling. Addresses quadratic attention cost bottleneck.SourceSource pageHuggingFace Daily Papers↳ Follow the threadUpdate thread / Follow-up threadNeoMME gets within 0.002 nDCG of ColQwen2.5 on document retrieval using 14x fewer parametersHugging Face BlogPolicy dependency / Stack layerLLMs Make More and Larger Edits on Code Written by a Different LLMarXiv 2609.03894Stack layer / ContrastThe Post-Training Method, Not the Data, Decides How Refusal Is Computed Inside a ModelarXiv 2609.03887Stack layer / ContrastCROSS-CATEGORY: Three Unrelated Orgs Shipped Coding-Agent Cost Routing in the Same 48 HoursGitHub Blog, Spotify Engineering and CodeRabbitStack layer / Threat patternA Blockchain-Anchored Black Box for Agent Workflows, Explicitly Scoped to Evidence Rather Than PreventionarXiv 2609.04017Stack layer / Update threadLLaDA-Image Sets Open-Source SOTA on Qwen-Image-Bench With a 6B DiT and Fully Released Training RecipesarXiv 2609.03796Stack layer / ContrastSemantic Similarity Misses the Model Diversity That Actually Predicts Correlated FailurearXiv 2609.03422Contrast / Update threadSide-Channel Papers Still Benchmark Against Proxies Set by Early Attack Work, Including at Top-Tier VenuesarXiv 2609.03893