Hacker News
IEEE Spectrum Maps the 2026 Inference Silicon Field: Etched Claims 500,000 Tokens/Sec, Majestic Labs 128TB of DRAM Per Rack
Matthew S. Smith's September 15 survey puts numbers to the post-GPU inference startups alongside the incumbents. Etched's Sohu transformer ASIC claims 500,000 tokens per second, Tensordyne's Napier up to 1,300 tokens per second per user, and a Cerebras WSE-3 deployment over 1,000 tokens per second off 44GB of SRAM holding 40-80B parameters. Majestic Labs' memory-aggregation rack claims up to 128TB of DRAM against roughly 20TB of HBM3E in NVIDIA's GB300 NVL72, and the piece notes a DeepSeek-R1 quantization that tripled throughput at under 1% benchmark loss.
Source
↳ Follow the thread