r/LocalLLaMA Debates China's DFSX Claim of 960 TB/s vs GB200 NVL72's 576 TB/s — on a 14nm Node
r/LocalLLaMA (432 upvotes / 152 comments), corroborated by Wccftech and TrendForce·high signal
A r/LocalLLaMA thread hit 432 upvotes and 152 comments on DFSX (Dongfang Suanxin) claiming a TY64 SuperNode built from 14nm DF2000 chips delivers 960 TB/s of memory bandwidth against 576 TB/s for NVIDIA's GB200 NVL72, achieved by stacking DRAM directly on logic at the wafer level instead of pairing external HBM. The tradeoff is stark and under-discussed in the thread: the same TY64 node offers 64 PFLOPS of BF16 compute versus GB200 NVL72's 360 PFLOPS. For anyone running memory-bound inference rather than training, that ratio is the whole argument — DF2000 targets Q4 2026 at 1000 TFLOPS BF16 and 15 TB/s per chip.