Tools
A single DGX Spark now serves DeepSeek V4 Flash at 3.0bpw with a 439,622-token KV pool, no second node
MiaAI-Lab/DeepSeek-v4-Flash-One-DGX-Spark (created 2026-08-20, MIT, 116 stars and 15 forks) is a Docker recipe serving the 3.0 bpw EXL3 build of DeepSeek V4 Flash 0731 on one GB10 DGX Spark at tensor-parallel 1, where the official FP4 build needs TP2 across two Sparks. Defaults changed on 2026-08-21 to a deep-context NVFP4 config: native 432-byte KV records, 0.94 GPU memory utilization, 384,000 max model length, DSpark K5 speculative decoding with a K64 draft, yielding a 439,622-token KV pool and a 370,104-token context stress-tested with exact needle recall. Weights are ~107 GB and download locally by default.
Source
↳ Follow the thread