Fetching from the wire…
Infra2026-09-16 · source-backed
PR #56935 fuses Q RoPE, sparse attention, the output's inverse RoPE and its FP8 cast into one launch writing straight into the buffer wo_a consumes, and this revision makes it the default rather than opt-in. It adds an nvfp4_ds_mla KV format: a 288-byte compressed record of 256 e2m1 pairs plus 32 e4m3 scales over 16 dims. Two orthogonal knobs ride the existing fused insert op, pinned by a bit-for-bit test over all eight combinations of 16/64 heads by norm by RoPE.
Each link below shares sources, entities, or timing with this story.
RTK has almost 80,000 GitHub stars and a simple promise. It sits between your coding agent and the shell, trims noisy command output before the model reads it, and claims 60-90% savings. Quesma ran it on Terminal-Bench 2.1 and found costs went up. With RTK on, average cost per...
PR #56893, merged September 15, switches V4.1 off the V4 paged fp8_ds_mla record onto a V4.1-specific one that quantizes the RoPE dims too: 512 B of fp8 e4m3 plus 16 UE8M0 scales, one per 32 dims, against V4's 448 B fp8 NoPE plus 128 B bf16 RoPE and 7 scales (GitHub). Because...
DeepSeek released V4 on April 24 and the numbers demand attention. V4-Pro is 1.6 trillion parameters total with 49 billion active, MIT-licensed, native 1M-token context. It scores 80.6% on SWE-bench Verified, putting it within 0.2 points of Claude Opus 4.6. On Terminal-Bench 2...
Sebastian Raschka's July 28 teardown argues K3 is less exotic than the release framing suggests: a scaled production version of Kimi Linear with Kimi Delta Attention as the hybrid attention layer and LatentMoE compressing large linear layers by down-projection. The genuinely n...
DeepSelect is the TopK kernel behind the indexer in DeepSeek Sparse Attention, and DeepSeek says it runs 2-20x faster than torch.topk. DeepJIT is a header-only C++20 runtime that compiles kernels at runtime, with one interface for both NVIDIA CUDA and Huawei Ascend. deepseek-r...
The file is 1.56TB. That's the first thing you notice about the moonshotai/Kimi-K3 Hugging Face repo that went live today. 2.8 trillion total parameters, 104B activated, 896 routed experts with 16 selected plus 2 shared per token, 93 layers split 69 Kimi Delta Attention and 24...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.