Fetching from the wire…
Infra2026-08-09 · source-backed
Using GB200 NVL72 disaggregated serving, 4–8 prefill endpoints at DEP2 plus one decode endpoint at DEP8, 8192-in/1024-out (vLLM). Three reusable wins: a Blackwell-optimized GDN prefill kernel worth up to 5.92x kernel and 1.13x prefill throughput via --gdn-prefill-backend flashinfer; a hybrid cache/state transfer moving both attention KV and SSM state correctly, cutting descriptors from 4,284 to 1,650 for ~7% throughput; and two race-condition fixes that finally made --async-scheduling viable. Full recipe includes VLLM_SSM_CONV_STATE_LAYOUT=DS, --mamba-ssm-cache-dtype bfloat16, --language-model-only, and --max-num-batched-tokens 16384 at 2x input sequence length.
Each link below shares sources, entities, or timing with this story.
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
The comparison is against GB300 NVL72, with 35x lower cost per million tokens, measured on the SemiAnalysis AgentX benchmark using real recorded agentic coding sessions with context growth, tool calls and sub-agent spawning preserved (NVIDIA). DeepSeek V4 Pro and Qwen3.5 were...
Nvidia set new MLPerf Inference v6.0 records on April 2 using four GB300 NVL72 systems (288 Blackwell Ultra GPUs) interconnected via Quantum-X800 InfiniBand. The headline number: 2.49 million tokens per second on DeepSeek-R1 in offline mode. That's the largest GPU configuratio...
Staged on ModelScope for 23:00 Beijing time August 26, it's roughly 125B parameters plus a separate N-gram embedding table of about 51B, activating 6B per token, with GDN gated-delta hybrid layers and Qwen Sparse Attention. Alibaba frames it as a technology preview of the Qwen...
Six co-designed chips, supply chain twice the size of Grace Blackwell, with AWS, Google Cloud, Microsoft, and OCI deploying instances in H2 2026 (NVIDIA). If inference really drops 10x, the economics of always-on agents change at the root. The cost crisis in story one is partl...
The day-0 SGLang post dated August 26 gives the architecture the release megathread didn't: 125B main parameters plus a separate 51.2B Per-Layer Embedding table at about 95.4 GiB in BF16, 6B activated per token, 48 layers split 36 GDN linear attention and 12 QSA sparse attenti...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.