Fresh llama.cpp Bugs Cut DeepSeek V4 Flash Throughput 6.4x on ROCm and Corrupt Quantized KV Cache
llama.cpp GitHub·medium signal
Two eval bugs filed in the last 48 hours limit DeepSeek V4 Flash as a local coding-agent backend. Issue #26399 (2026-08-01) reports GGML_OP_TOP_K falling back to CPU on HIP/ROCm above roughly 3-4K context, costing a 6.4x token-generation loss exactly where agent contexts live. Issue #26423 (2026-08-02) reports quantized KV cache still producing garbage on master even with #25202, and points at the arch-scoped rotation disable from #25391 as the fix. Both are open.