IEEE Spectrum maps the 2026 inference silicon fight: GPUs idle 50-80% of the time waiting on memory
The 2026-09-15 feature frames inference as a memory-bandwidth problem rather than a compute one and catalogs the contenders: Nvidia's $20B Groq acquisition producing the Groq 3 LPU with 500MB on-chip SRAM and 7x GPU memory bandwidth, paired with Vera Rubin for two-phase inference; AWS with Cerebras WSE-3 at 44GB on-chip SRAM, which OpenAI used to push GPT-5.3-Codex-Spark past 1,000 tokens/second against 50-125 for standard GPT-5.4. Startups split along a memory axis — d-Matrix stacking logic on DRAM, Majestic Labs at 128TB DRAM per rack versus GB300's ~20TB HBM3E, Tensordyne's Napier at 1,300 tokens/sec on logarithmic number systems, Etched's Sohu at 500,000 tokens/sec for Llama 70B. The cost driver behind the split is simple: HBM runs two to three times the price of commodity DRAM.
Source
↳ Follow the thread