Skills
EAGLE-3 speculative decoding on vLLM is delivering ~2x throughput on AMD Instinct
As of July 13, 2026, AMD Quark is training, quantizing, and serving EAGLE-3 speculative-decoding drafts through vLLM on AMD Instinct GPUs, reporting up to 2.00x throughput for Kimi-K2.5 and 1.79x for MiniMax-M2.5. Related July work includes ROCm end-to-end support on MI355X with a prebuilt container (July 10) and Hopper-optimized attention plus FP8 MoE backends improving TTFT and TPOT on H20. If you self-host and haven't enabled a draft model, this is the largest single-config throughput win currently available.
Source
↳ Follow the thread