Reddit
A Hand-Optimized Qwen3.8-27B Stack Hit 381 Tokens/Sec Single-Request on One RTX 3090, Up From 82 Four Days Ago
The author of a hyper-optimized Qwen3.8-27B inference engine for the RTX 3090 posted a follow-up: 82 tps single-request four days ago, then ~114, then ~138 with DFlash2 drafting and lookup-augmented drafting, and now ~133 tps on real chat prompts with 382 tps when the model reproduces its own context. The stack is fp8 KV cache, int8 lm_head and embed_tokens, fp16 recurrent state, int8 activations, W4A16-requantized DFlash2 block drafting, prefix caching, split-KV verify attention and KVarN for 262k context. The number they care about is 15 of 16 tokens accepted per verify step on document-quoting workloads.
Source
↳ Follow the thread