vLLM ROCm Added to Lemonade as Experimental Backend — Self-Contained Bundle, No System ROCm Required
r/LocalLLaMA / Lemonade·medium signal
Lemonade announced vLLM as an experimental AMD ROCm backend on May 8, enabling day-0 model support directly from Hugging Face safetensors without GGUF conversion. The backend ships as a self-contained bundle with relocatable Python, PyTorch ROCm, and all ROCm libraries — zero system dependencies. The r/LocalLLaMA post (317 upvotes, 69 comments) highlights paged-attention KV cache, continuous batching, chunked prefill, and tensor parallelism as key capabilities for local serving workloads.