NVIDIA/Model-Optimizer: Unified SOTA Model Compression Library for Quantization, Pruning, and Speculative Decoding at 2.3K Stars
GitHub·low signal
NVIDIA's Model-Optimizer unifies quantization, pruning, distillation, and speculative decoding into a single library that compresses models for deployment to TensorRT-LLM, TensorRT, and vLLM. Following NVIDIA's GTC 2026 push into open-source agent tooling (including the new OpenShell runtime), this library represents the production optimization layer that bridges the gap between model training and efficient deployment at scale.