Fetching from the wire…
Models2026-09-14 · source-backed
The listing gives 84GB ECC GDDR7, 1,398 GB/s bandwidth, 21,760 CUDA cores, PCIe Gen 5 x16, up to 600W, positioned between the PRO 6000 and 5000. Roughly triple a 5090's VRAM in one slot-class card. No pricing published and the page reads "Coming Soon"; the RTX PRO 6000 Blackwell repriced to about $16,000 in August is the anchor.
Each link below shares sources, entities, or timing with this story.
NVIDIA's Blackwell successor is in production ahead of schedule. The NVL72 rack (72 GPUs) delivers 3.6 exaFLOPS for inference, with 288GB HBM4 per GPU. NVIDIA claims 10x lower cost-per-token versus Blackwell. The Rubin CPX variant — purpose-built for million-token inference —...
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
NVIDIA claims 1.8x faster task completion and twice the efficiency against traditional x86, with Vera Rubin NVL72 racking 72 Rubin GPUs and 36 Vera CPUs alongside ConnectX-9 SuperNICs and BlueField-4 DPUs. (NVIDIA) Architecture detail months before shipping, two days ahead of...
Portable Computer launched August 26, running the orchestrator LLM, subagent LLM, planner, tool router, scheduler and local search index locally, with local work consuming no billing credits and each cloud escalation requiring separate approval (VentureBeat). Launch platform i...
TechCrunch's August 29 piece frames Nvidia's durable advantage as system-level, built around Vera Rubin pairing the Rubin GPU with the Vera CPU, a Groq 3 LPX inference accelerator, and storage and networking racks. VP of storage technology Jason Hardy is quoted claiming "upwar...
June local-inference benchmarks across the 128GB class put NVIDIA's DGX Spark (~$4k), AMD's Strix Halo / Ryzen AI Max+ 395 (~$2 to 3k), and the M5 Max 128GB (~$5k) head to head (Hardware Corner). Prompt processing favors CUDA hard. But token generation lands at a surprisingly...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.