Reddit
An expert-cache PR lifts Qwen3.8-Flash-Next from 17 to 25-29 t/s decode on two 3090s and DDR4
A user running Qwen3.8-Flash-Next at unsloth UD-Q6_K_XL on 2x RTX 3090 over PCIe 3.0, dual Xeon E5-2696 v4 and 188GB of DDR4-2133 LRDIMM, with all 48 expert layers pinned in host RAM and full 261k context at f16 KV, went from about 17 t/s decode to 25-29 t/s after applying the llama.cpp expert cache PR. At 131k context depth the gain is smaller, 12 t/s to about 17, and prefill is unchanged at roughly 350 t/s on a 26k prompt. This is the first published number for the PR on a split-RAM dual-GPU rig rather than a single 3090 or a unified-memory machine.
Source
↳ Follow the thread