antirez's ds4 Adds DeepSeek V4 PRO Support and CUDA/ROCm Backends — 19,834 Stars, +150 Today, 26 tok/s on a 128GB M3 Max
GitHub·high signal
Salvatore Sanfilippo's single-file C inference engine for DeepSeek V4 Flash has widened from Metal-only to Metal, CUDA and ROCm, and now covers V4 PRO as well as Flash, with a last push on 2026-08-01. It is deliberately narrow — not a GGUF runner, not a wrapper — shipping DS4-specific model loading, KV cache handling, prompt rendering and an OpenAI-compatible server. Third-party reviews report the 284B MoE running at 26 tokens/sec at 50W peak draw on a MacBook Pro M3 Max with 128GB under antirez's own Q2 quantization.