Fetching from the wire…
Infra2026-09-17 · source-backed
NVIDIA posted v6.1 results on September 16 across DeepSeek-R1, Qwen3-VL, WAN 2.2 text-to-video, GPT-OSS-120B, DLRMv3 and an edge-agentic Qwen3.6-27B test. Vera Rubin NVL72 delivered up to 3.7x GB300 NVL72 throughput on Qwen3-VL and up to 2.5x on DeepSeek-R1. GB300 NVL72 separately scaled from 72 to 288 GPUs at 99% efficiency on DeepSeek-R1 offline, and software alone (kernel fusion, disaggregated serving) moved GB300's Qwen3-VL numbers up to 1.6x between v6.0 and v6.1. That software delta is the one I'd watch, because it applies to hardware you may already own.
Each link below shares sources, entities, or timing with this story.
Nvidia set new MLPerf Inference v6.0 records on April 2 using four GB300 NVL72 systems (288 Blackwell Ultra GPUs) interconnected via Quantum-X800 InfiniBand. The headline number: 2.49 million tokens per second on DeepSeek-R1 in offline mode. That's the largest GPU configuratio...
The September 15 feature catalogs GPUs idling 50-80% of the time waiting on memory, then maps the contenders: Nvidia's $20B Groq acquisition producing an LPU with 500MB on-chip SRAM and 7x GPU memory bandwidth, a Cerebras WSE-3 deployment pushing GPT-5.3-Codex-Spark past 1,000...
The comparison is against GB300 NVL72, with 35x lower cost per million tokens, measured on the SemiAnalysis AgentX benchmark using real recorded agentic coding sessions with context growth, tool calls and sub-agent spawning preserved (NVIDIA). DeepSeek V4 Pro and Qwen3.5 were...
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
NVIDIA claims 1.8x faster task completion and twice the efficiency against traditional x86, with Vera Rubin NVL72 racking 72 Rubin GPUs and 36 Vera CPUs alongside ConnectX-9 SuperNICs and BlueField-4 DPUs. (NVIDIA) Architecture detail months before shipping, two days ahead of...
Salesforce released Koa, built by post-training Nemotron-3-Super-120B, Nvidia's open-weight hybrid Mamba-Transformer MoE with 120B total and 12B active parameters (TechCrunch, paper at arXiv 2609.15066). The training was GRPO reinforcement learning on public and synthetic data...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.