News
d-Matrix Raptor Claims 100 TB/s From 3D-DRAM and 20x the Bandwidth Density of HBM
d-Matrix presented Raptor at Hot Chips 2026, a 3D-DRAM inference accelerator with 32GB per card reaching roughly 100 TB/s, and claims 32.6 GB/s per mm2 against about 1.5 GB/s for HBM parts, plus 2.96 mW per GB/s versus 40 mW, a 13.5x power-efficiency gain. A 72-card rack is pitched as holding a frontier model at 1M context using 4-bit weights and an 8-bit KV cache, with roughly 1,000 tokens per second per user on a 3-trillion-parameter model. No ship date or commercial availability was given, so treat the numbers as prototype claims against NVIDIA's Rubin R200 and HBM4.
Source
↳ Follow the thread