Fetching from the wire…
Infra2026-08-20 · source-backed
Two published vLLM configs for a single 24GB card at 250W with 150k context: batch mode measuring ~1,094 tok/s steady-state decode at 64 concurrent (942 end-to-end, rising to ~1,222/1,042 with all layers int8), and single-user mode at 114-122 tok/s single-stream via MTP speculation with four cheap drafts, calibrated int4 lm_head and split-KV verify attention. (GitHub) The standout number: 381 tok/s at 25k context when the model reproduces its own context, quoting a document or applying an edit, using DFlash2 with 15 drafted tokens per verify step. Speculation wins below roughly 8 concurrent users, plain batching wins above.
Each link below shares sources, entities, or timing with this story.
Shared entities / Earlier coverage / Tension
Both cover GitHub, Qwen3, RTX; earlier GitHub coverage from 2026-04-23; pushes against this story (versus).
Shared entities / Same source domain / Earlier coverage / Tension
Both cover GitHub, Qwen3; reported by the same outlet (github.com); earlier GitHub coverage from 2026-08-17.
Shared entities / Tension
Both cover MTP, Qwen3, RTX; pushes against this story (against).
vLLM supports MTP / Shared entity: GitHub / Same source domain / Earlier coverage
Linked by a graph relationship (vLLM supports MTP); both cover GitHub; reported by the same outlet (github.com).
Shared entity: Qwen3 / Same source domain / Shared topic / Earlier coverage / Downstream implication
Both cover Qwen3; reported by the same outlet (github.com); overlapping topics (card, context).
Shared entities / Shared topic / Earlier coverage
Both cover MTP, Qwen3; overlapping topics (card, context); earlier MTP coverage from 2026-08-18.
Shared entities / Same source domain / Earlier coverage
Both cover GitHub, Qwen3; reported by the same outlet (github.com); earlier GitHub coverage from 2026-08-17.
Both cover GitHub, Qwen3; reported by the same outlet (github.com); earlier GitHub coverage from 2026-05-04.