Fetching from the wire…
Infra2026-08-27 · source-backed
The day-0 SGLang post dated August 26 gives the architecture the release megathread didn't: 125B main parameters plus a separate 51.2B Per-Layer Embedding table at about 95.4 GiB in BF16, 6B activated per token, 48 layers split 36 GDN linear attention and 12 QSA sparse attention, 512 experts with top-10 routing (LMSYS). Sparse pinned-host offload of the PLE table on H200 with TP4 drops weight footprint per GPU from 83.91 to 60.45 GiB and raises KV cache capacity from 1.84M to 3.28M tokens, a 78.5% increase. On B200 TP4 with MTP speculative decoding it reaches 540 tok/s at batch size 1 with an accept length of 3.3.
Each link below shares sources, entities, or timing with this story.
Shared entities / Shared topic / Earlier coverage
Both cover August, Flash, GDN, Next; overlapping topics (architecture, attention, august, cache, sparse); earlier August coverage from 2026-08-26.
Shared entities / Shared topic
Both cover August, Flash, MTP, Next; overlapping topics (architecture, attention, august, cache).
Shared entities / Same source domain / Shared topic / Earlier coverage
Both cover Flash, GiB, H200, LMSYS; reported by the same outlet (lmsys.org); overlapping topics (attention, cache, token).
Google released Flash / Shared entities / Earlier coverage
Linked by a graph relationship (Google released Flash); both cover August, Flash, Qwen3; earlier August coverage from 2026-08-07.
Shared entities / Shared topic / Earlier coverage
Both cover August, BF16, GPU; overlapping topics (attention, august, bf16); earlier August coverage from 2026-08-11.
Both cover GPU, MTP, SGLang; overlapping topics (attention, cache, decoding); earlier GPU coverage from 2026-07-25.
Shared entities / Earlier coverage
Both cover August, Flash, Qwen3, SGLang; earlier August coverage from 2026-08-10.
Shared entities / Shared topic / Earlier coverage
Both cover GDN, GPU, Qwen3; overlapping topics (attention, cache); earlier GDN coverage from 2026-08-09.