Reddit
MXFP4 with W4A8 kernels pushes Qwen3.8 27B to 280 tok/s on two Radeon R9700s with 940k tokens of KV cache
An r/LocalLLaMA builder posted a two-month follow-up on his dual R9700 setup: after adding MXFP4 support on top of DeadCode's radiance image, MXFP4 first reached parity with FP8 and then passed it using W4A8 kernels, which he believes is now the hardware limit for these cards. BetterBench decode results for Qwen3.8 27B with DFlash2 range from 280.0 t/s on JSON down to 116.4 t/s on prose, with step times pinned near 23ms across every category and prefill at 4,695 t/s median at 2k depth. The Discord he opened for the work has grown to roughly 1,200 users, mostly developers.
Source
↳ Follow the thread