Skills
Google Splits TPU into Two Chips for the Agentic Era: TPU 8t for Training (9,600-Chip Clusters, 2PB HBM) and TPU 8i for Inference (3x SRAM, 80% Better Price-Performance)
Google announced its first-ever split architecture: TPU 8t optimized for training scales to 9,600 chips with 2 Petabytes of shared HBM and nearly 3x compute over Ironwood, cutting frontier model development from months to weeks. TPU 8i for inference features 288GB HBM and 384MB on-chip SRAM (3x predecessor), keeping entire model working sets on-chip to reduce latency for multi-agent tasks. Google claims 80% price-performance improvement for 8i and 2.8x gain for 8t. The split signals that one-size-fits-all accelerators are dead for the agent era.
Source
↳ Follow the thread