r/LocalLLaMA Argues DeepSeek V4 Flash 0731 Is the 'Killer App' That Sells DGX Sparks — 26 tok/s on One Box, 82 on Two
A 216-upvote post that pulled 233 comments — a comment-heavy 1.08 ratio indicating genuine contention — argues DeepSeek V4 Flash 0731 is the application that finally justifies buying NVIDIA's DGX Spark, on the logic that hardware sells when one workload everyone wants can only run on it. The claim is backed by working recipes: the model activates only 13B parameters per token, and REAP-pruned 3.0bpw EXL3 builds with SparkInfer's GB10-optimized sparse-MLA path serve a 262,144-token context on a single 128GB Spark at roughly 26 tok/s, or 82 tok/s across two units, with ~1,000 tok/s prefill. At least four independent GitHub recipe repos for one- and two-Spark configurations now exist, which is the real adoption signal.
Source
↳ Follow the thread