AMD's Lemonade server drops the OpenMOSS ROCm backend after measuring it about 40x slower than Vulkan
Phoronix (via r/LocalLLaMA)·low signal
Lemonade 2026.40 RC can now stream models on AMD APUs such as Strix Halo by sizing against the GTT pool instead of the fixed vRAM carve-out, which fixes DeepSeek-V4-Flash failures. It removes OpenMOSS ROCm on Windows and Linux and adds a launch agent for JetBrains Junie. Anyone running local inference on AMD should use the Vulkan backend.