Fetching from the wire…
Research2026-09-23 · source-backed
arXiv 2609.25624 uses fixed-configuration fused-upcast GEMMs: load 16-bit weights, accumulate in FP32, use a reduction order depending only on problem shape. Every GPU runs the same operation sequence. 1.17-3.1x faster end-to-end than the previous deterministic approach, with half the weight-memory traffic. This is the one to reach for if you run replayable evals across a mixed GPU fleet.
Each link below shares sources, entities, or timing with this story.
NVIDIA and AWS announced June 23 that NVIDIA's cuVS library now powers GPU-accelerated vector indexing as the default in Amazon OpenSearch Serverless, claiming up to 10x faster index builds at roughly a quarter the cost versus CPU-only, making billion-scale vector DBs buildabl...
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
NVIDIA's Blackwell successor is in production ahead of schedule. The NVL72 rack (72 GPUs) delivers 3.6 exaFLOPS for inference, with 288GB HBM4 per GPU. NVIDIA claims 10x lower cost-per-token versus Blackwell. The Rubin CPX variant — purpose-built for million-token inference —...
The file is 1.56TB. That's the first thing you notice about the moonshotai/Kimi-K3 Hugging Face repo that went live today. 2.8 trillion total parameters, 104B activated, 896 routed experts with 16 selected plus 2 shared per token, 93 layers split 69 Kimi Delta Attention and 24...
Nvidia set new MLPerf Inference v6.0 records on April 2 using four GB300 NVL72 systems (288 Blackwell Ultra GPUs) interconnected via Quantum-X800 InfiniBand. The headline number: 2.49 million tokens per second on DeepSeek-R1 in offline mode. That's the largest GPU configuratio...
A rare end-to-end systems report for trillion-parameter MoE post-training outside the GPU world: hierarchical optimization across model parallelism, computation-communication orchestration, and low-level kernels on an Ascend NPU SuperPOD, a 2.93× improvement over the open-sour...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.