Reddit
Micron's Hot Chips 2026 memory-wall chart, and the Samsung processing-in-memory number underneath it
A chart from Raghu Sreeramaneni's memory tutorial at Hot Chips 2026 plots normalized TFLOPs from TPU v3 through R200 against HBM2e through HBM4, both log scale: compute roughly 3x every two years, HBM bandwidth under 2x. The three mitigations on the slide are memory beside compute, memory on a shorter link, and multiply units inside the memory itself. The concrete datapoint most commenters latched onto is Samsung shipping that third approach in LPDDR5X with a measured 3.01x tokens per second on Llama 3.1 8B, which attacks the wall rather than moving it.
↳ Follow the thread