Tools
llama.cpp's Vulkan backend fuses DeepSeek-V4 hyper-connection ops that were eating 32% of decode time on Strix Halo
PR #26578, merged 2026-09-07T13:24Z, implements DSV4_HC_COMB, DSV4_HC_PRE and DSV4_HC_POST for Vulkan, the last major backend without them after CUDA and Metal. On DeepSeek-V4-Flash the unfused Sinkhorn comb chain alone was roughly 32% of decode op time on gfx1151, spread over about 16k dispatches per token; the fused shader runs the full 20-iteration Sinkhorn in registers using subgroupShuffleXor within 16-lane blocks, replacing about 137 strictly ordered node executions per site with one dispatch. Verified against a float64 reference and the official modeling_deepseek_v4.py to about 6e-7.
Source
↳ Follow the thread