Fetching from the wire…
Source-backed findings, relationship evidence, citations, and briefing history from the public MindPattern archive.
Showing the first 40 findings. More graph evidence exists in the corpus.
Compiled 2026-07-21 · source-backed
NetraRuntime kernel optimizations achieved 78,498 output tokens per second, 2.16x vLLM throughput.
Source findingmarin integrates vLLM in its training infrastructure.
Source findingvToken integrates with vLLM with token-table indirection.
Source findingvLLM shipped day-0 support for Qwen3.8-2.4T-A95B on NVIDIA hardware.
Source findingvllm.cpp demonstrated token-identical output with vLLM at equal or faster speeds
Source findingvLLM shipped Decode Context Parallelism (DCP) for long-context KV cache optimization.
Source findingvLLM achieved 25K tok/s per GB200 GPU using FlashInfer backend.
Source findingvLLM uses FlashInfer GDN backend for prefill kernel optimization.
Source findingvLLM shipped day-0 support for Kimi K3.
Source findingT3MP3ST supports offline operation via vLLM
Source findingTransformers-defined models now run inside vLLM at native speed.
Source findingXGrammar is default structured output backend in vLLM 0.4+
Source finding