181 tok/s Aggregate on 2x DGX Spark, With a mmap Fix That Cut One Prefill's Disk Reads From 603 GB to 19 GB
A builder posted a full recipe hitting 181 tok/s aggregate (peaking at 195) across ~9 concurrent agent sessions on two DGX Sparks running Qwen3.8-Flash-Next NVFP4 at 512K context, with single-stream decode at 30-50 tok/s. The nodes are joined by a direct ConnectX-7 cable over NCCL/RoCE at 200 Gb with TP=2, and the post warns the TCP fallback is silent and costs half the speed, so you must confirm 'Using network IB' in the NCCL log. Mapping the 47.7 GiB FP8 n-gram table off NVMe dropped per-node weights from 65 to 41 GiB, and two fixes made it fast: madvise(MADV_RANDOM) killed a 30x read amplification (one 405K prefill read 603 GB from disk before, 19 GB after) and 64 gather threads removed fault-latency serialization, freeing memory for a 2.89M-token KV pool.
Source
↳ Follow the thread