OSS
llama.cpp Ships b8825 on April 17 — Four Tagged Releases in Five Days Maintain Breakneck Inference Engine Cadence
llama.cpp pushed release b8825 on April 17, continuing a cadence of four tagged releases between April 12-17. Recent improvements include Vulkan flash attention with DP4A shader for quantized KV cache, dot product precision fixes for the Vulkan backend, and Wave32 scalar flash attention refactor for AMD GPUs. The project remains the backbone of local LLM inference for hundreds of thousands of developers, with each release expanding hardware compatibility.
Source
↳ Follow the thread