llama.cpp's Vulkan backend gets an int8 cooperative-matrix matmul path for AMD RDNA3 and RDNA4
GitHub·medium signal
llama.cpp PR #27952 (merged 2026-09-24) adds an int8 coopmat1 MMQ shader, so Vulkan no longer has to dequantize to fp16 before using matrix cores on AMD RDNA3/RDNA4. It covers q4_0 through q6_k, mxfp4, nvfp4 and iq4_nl. On a Strix Halo Radeon 8060S, q4_0 MUL_MAT improves 1.29x over master and also beats ROCm. The author reports RDNA4 as roughly neutral except for MoE prompt processing, and slower quants are disabled there.