Reddit
Lemonade Now Fronts 15 Inference Engines Behind One Base URL With Semantic and Policy Routing
The Lemonade maintainers posted an end-of-summer update on 2026-08-26 covering CUDA, ARM64, Metal and Vulkan backends for all core engines, experimental engines for music and 3D asset generation, and a router that now does semantic and policy routing. It installs as a single OS service managing models and engines behind one base URL, and ships as an embeddable SDK. Maintainers confirmed in-thread that DGX Spark is supported, and noted you can point the config at your own llama-server binary to run llama.cpp PRs early while keeping Lemonade's routing.
↳ Follow the thread