An AI Chip Survey Puts TPU v8t at 121 ExaFLOPS FP4 Across 9,600 Chips and Notes NVIDIA Bought Groq for About $20B
Jacob Peake's essay compares NVIDIA (Volta through Rubin Ultra), Google TPU (v1 through v8), AMD CDNA, Cerebras WSE and Groq's LPU, with a specification table listing B200 at 1.457 ExaFLOPS FP4 and 192GB HBM3e, MI355X at 2.9 ExaFLOPS FP4 with 288GB HBM3E and 185B transistors, and WSE-3 at 125 PFLOPS sparse FP16 across a 46,225 mm² wafer with 900K cores and 44GB of SRAM. The thesis is that the memory wall forces inverse tradeoffs, NVIDIA hiding latency with multithreading, Google using compiler-scheduled static dataflow with no caches or scheduler, AMD going memory-first with a large last-level cache, and Cerebras matching compute to bandwidth by deleting the die boundary. It also notes FLOPs halve each generation from FP16 to FP8 to FP4, with block-scale quantization restoring accuracy.
↳ Follow the thread