Fetching from the wire…
Public story · 2026-08-29 · source-backed
A Hugging Face repo packages the model with the n-gram lookup table offloaded to SSD and streamed. The credible reply in the thread: a builder on an RTX Pro 6000 running the RAM variant reports over 12k prefill and over 170 tok/s single-stream decode, plus 440 tok/s at concurrency 4 on a 500W power-limited workstation. Another runs the SSD version on a 5090 with 64GB DDR5 at 34 tps and 55 pp, and swapped the q4 n-gram table for bf16 with no speed penalty and better output. The reason it works, per the thread, is that n-gram lookups are predictable enough to prefetch, so SSD latency stops mattering. (r/LocalLLaMA)
Each link below shares sources, entities, or timing with this story.
Hugging Face criticizes OpenAI / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (Hugging Face criticizes OpenAI); both cover Flash, Hugging Face, LocalLLaMA, Qwen3; reported by the same outlet (reddit.com).
Hugging Face released Safetensors / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Hugging Face released Safetensors); both cover Hugging Face, LocalLLaMA; reported by the same outlet (reddit.com).
Hugging Face criticizes OpenAI / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Hugging Face criticizes OpenAI); both cover Hugging Face, LocalLLaMA; overlapping topics (face, hugging).
Hugging Face criticizes OpenAI / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (Hugging Face criticizes OpenAI); both cover Flash, Hugging Face, LocalLLaMA; reported by the same outlet (reddit.com).
Hugging Face partners with NVIDIA / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Hugging Face partners with NVIDIA); both cover Hugging Face, LocalLLaMA; reported by the same outlet (reddit.com).
Shared entities / Shared topic / Earlier coverage
Both cover Flash, Next, Qwen3, RAM; overlapping topics (flash-next, n-gram, table); earlier Flash coverage from 2026-08-28.
Hugging Face released Skills / Shared entities / Earlier coverage
Linked by a graph relationship (Hugging Face released Skills); both cover Flash, Hugging Face; earlier Flash coverage from 2026-05-20.
Hugging Face partners with NVIDIA / Shared entities / Same source domain / Earlier coverage / Downstream implication
Linked by a graph relationship (Hugging Face partners with NVIDIA); both cover LocalLLaMA, Qwen3, RTX PRO; reported by the same outlet (reddit.com).