vLLM v0.19.0 (April 2) ships complete Gemma 4 support for all four model variants, handling MoE routing, multimodal inputs, reasoning traces, and tool-use natively. The async scheduler — overlapping engine scheduling with GPU execution — is now on by default with zero configuration. At 78K GitHub stars and trending, vLLM remains the dominant open-source inference engine for production LLM serving.