Fetching from the wire…
Infra2026-09-04 · source-backed
Blackwell's FP4 tensor cores don't automatically speed up attention, because softmax conversion and on-chip dependencies dominate once the matrix products shrink. Direct-P maps scores directly to FP4 probabilities for noncausal inference. A separate causal path reconstructs probabilities from saved quantized queries and keys and uses FP8 gradient operands, accelerating a full single-GPU 8B update by up to 1.14x. The caveat is sharp: matched distributed training must retain FP8 probabilities and values, because every tested MXFP4 probability/value training trajectory diverged. arXiv 2609.04105
Each link below shares sources, entities, or timing with this story.
NVIDIA's Blackwell successor is in production ahead of schedule. The NVL72 rack (72 GPUs) delivers 3.6 exaFLOPS for inference, with 288GB HBM4 per GPU. NVIDIA claims 10x lower cost-per-token versus Blackwell. The Rubin CPX variant — purpose-built for million-token inference —...
TechCrunch's August 29 piece frames Nvidia's durable advantage as system-level, built around Vera Rubin pairing the Rubin GPU with the Vera CPU, a Groq 3 LPX inference accelerator, and storage and networking racks. VP of storage technology Jason Hardy is quoted claiming "upwar...
Six co-designed chips, supply chain twice the size of Grace Blackwell, with AWS, Google Cloud, Microsoft, and OCI deploying instances in H2 2026 (NVIDIA). If inference really drops 10x, the economics of always-on agents change at the root. The cost crisis in story one is partl...
This work pairs E2M1 payloads with unsigned E5M3 block scales whose wider range permits periodic tensor scaling, applies selective stochastic rounding only to backward gradients, drops the Hadamard transform entirely, and uses FP4 in every eligible internal linear. Pretraining...
The file is 1.56TB. That's the first thing you notice about the moonshotai/Kimi-K3 Hugging Face repo that went live today. 2.8 trillion total parameters, 104B activated, 896 routed experts with 16 selected plus 2 shared per token, 93 layers split 69 Kimi Delta Attention and 24...
The open-source inference engine released v0.20.2 with GPU-native Triton kernels, async speculative decoding with zero-bubble overlap, and GPU-less render serving for separating multimodal preprocessing from GPU inference. FP8 is now standard for H100 and Blackwell. If you're...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.