Fetching from the wire…
Infra2026-08-28 · source-backed
Build b10669 binds the f16 KV cache in place for the oneDNN SDPA path, and the commit does the arithmetic on Qwen3.8 27B Q4_K_S at a live KV length of 34,816: 71.3 MB per tensor, 142.6 MB staged per call for K and V, 285.2 MB per call, 16 calls per ubatch, so 4.56 GB of memory traffic for one prefill chunk (GitHub). It scales with live KV length, so the first ubatch at seq=2048 moves only 0.27 GB and the problem gets worse the longer you talk.
Each link below shares sources, entities, or timing with this story.
Shared entities / Same source domain / Shared topic / Earlier coverage
Both cover GitHub, Qwen3; reported by the same outlet (github.com); overlapping topics (call, llama, path).
Shared entities / Same source domain / Earlier coverage
Both cover Build, GitHub, Qwen3; reported by the same outlet (github.com); earlier Build coverage from 2026-02-17.
Shared entities / Same source domain / Shared topic / Earlier coverage
Both cover GitHub, Qwen3; reported by the same outlet (github.com); overlapping topics (memory, prefill).
Shared entities / Same source domain / Earlier coverage / Tension
Both cover GitHub, Qwen3; reported by the same outlet (github.com); earlier GitHub coverage from 2026-08-17.
Shared entities / Earlier coverage
Both cover Build, GitHub, Qwen3; earlier Build coverage from 2026-04-26.
Shared entity: GitHub / Same source domain / Shared topic / Earlier coverage
Both cover GitHub; reported by the same outlet (github.com); overlapping topics (cache, commit, memory).
Shared entity: Qwen3 / Same source domain / Shared topic / Earlier coverage
Both cover Qwen3; reported by the same outlet (github.com); overlapping topics (cache, llama, longer).
Shared entities / Same source domain / Earlier coverage
Both cover GitHub, Qwen3; reported by the same outlet (github.com); earlier GitHub coverage from 2026-08-27.