Fetching from the wire…
Infra2026-09-16 · source-backed
The September 15 feature catalogs GPUs idling 50-80% of the time waiting on memory, then maps the contenders: Nvidia's $20B Groq acquisition producing an LPU with 500MB on-chip SRAM and 7x GPU memory bandwidth, a Cerebras WSE-3 deployment pushing GPT-5.3-Codex-Spark past 1,000 tokens/second off 44GB of SRAM, Majestic Labs at 128TB DRAM per rack against GB300's ~20TB HBM3E, Etched's Sohu claiming 500,000 tokens/sec for Llama 70B. The cost driver behind the whole split is that HBM runs two to three times commodity DRAM.
Each link below shares sources, entities, or timing with this story.
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
OpenAI's Codex arrived on Windows (500K+ waitlist) with production-grade OS-level sandboxing (restricted tokens, filesystem ACLs, dedicated sandbox users) and a multi-agent management UI. GPT-5.3-Codex-Spark hits 1,000+ tokens/sec on Cerebras. Available across all ChatGPT tier...
TrendForce reporting via Tom's Hardware says prototype variants run 192GB or 256GB, down from the announced 288GB, with some configs using fewer than 16 stacks and substituting HBM4 for HBM4E. Driver is tightening HBM supply across SK hynix, Samsung, and Micron. At GTC 2025 co...
Adapted from ICLR 2026 and MLSS 2026 workshops, covering masked diffusion, block diffusion for variable-length generation, encoder-decoder architectures, remasking-based error correction, sampling distillation, guidance and RL post-training (Kuleshov Group). It catalogs what y...
TechCrunch's August 29 piece frames Nvidia's durable advantage as system-level, built around Vera Rubin pairing the Rubin GPU with the Vera CPU, a Groq 3 LPX inference accelerator, and storage and networking racks. VP of storage technology Jason Hardy is quoted claiming "upwar...
AMD unveiled its first rack-scale system to directly contest Nvidia at the rack level, with engineering samples in H2 2026 and mass production targeted Q2 2027. Microsoft joins Meta, OpenAI and Oracle as customers; Meta plans 1 gigawatt of Helios racks by year-end against a lo...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.