Fetching from the wire…
Models2026-08-21 · source-backed
The progression: 82, then ~114, then ~138 with DFlash2 drafting and lookup-augmented drafting, and now ~133 tok/s on real chat prompts with 382 tok/s when the model reproduces its own context. Stack is fp8 KV cache, int8 lm_head and embed_tokens, fp16 recurrent state, int8 activations, W4A16-requantized DFlash2 block drafting, prefix caching, split-KV verify attention, KVarN for 262k context. r/LocalLLaMA The number they care about is 15 of 16 tokens accepted per verify step on document-quoting workloads, which is the honest caveat: the 382 figure is a best case on a specific shape of work.
Each link below shares sources, entities, or timing with this story.
Shared entities / Shared topic / Earlier coverage
Both cover DFlash2, Qwen3, RTX; overlapping topics (attention, context, dflash2, verify); earlier DFlash2 coverage from 2026-08-20.
Shared entities / Same source domain / Earlier coverage / Tension
Both cover LocalLLaMA, Qwen3, RTX; reported by the same outlet (reddit.com); earlier LocalLLaMA coverage from 2026-08-20.
Both cover LocalLLaMA, Qwen3, RTX; reported by the same outlet (reddit.com); earlier LocalLLaMA coverage from 2026-04-23.
Shared entities / Same source domain / Shared topic / Earlier coverage
Both cover LocalLLaMA, Qwen3; reported by the same outlet (reddit.com); overlapping topics (attention, cache, context).
Shared entities / Same source domain / Earlier coverage
Both cover LocalLLaMA, Qwen3, RTX; reported by the same outlet (reddit.com); earlier LocalLLaMA coverage from 2026-08-10.
Shared entities / Same source domain / Shared topic / Earlier coverage
Both cover LocalLLaMA, Qwen3; reported by the same outlet (reddit.com); overlapping topics (cache, context).
Both cover LocalLLaMA, Qwen3; reported by the same outlet (reddit.com); overlapping topics (attention, context).
Both cover LocalLLaMA, RTX; reported by the same outlet (reddit.com); overlapping topics (cache, context).