Fetching from the wire…
Infra2026-04-21 · source-backed
kvcached now enables automatic prefix caching for vLLM and SGLang with elastic memory management. Multiple LLMs share a GPU without rigid partitioning. Red Hat's Sardeenz is building on it for dynamic multi-model Kubernetes serving.
Each link below shares sources, entities, or timing with this story.
NVIDIA supports Kubernetes / Shared entity: Red Hat / What happened next / Tension
Linked by a graph relationship (NVIDIA supports Kubernetes); both cover Red Hat; picks up the Red Hat thread on 2026-06-27.
OpenSRE uses Kubernetes / Shared entity: Kubernetes / Same source domain / What happened next
Linked by a graph relationship (OpenSRE uses Kubernetes); both cover Kubernetes; reported by the same outlet (github.com).
NVIDIA supports Kubernetes / Shared entity: GPU / What happened next / Tension
Linked by a graph relationship (NVIDIA supports Kubernetes); both cover GPU; picks up the GPU thread on 2026-06-18.
Linked by a graph relationship (NVIDIA supports Kubernetes); both cover GPU; picks up the GPU thread on 2026-06-25.
Sardeenz built by Red Hat / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (Sardeenz built by Red Hat); both cover GPU, Red Hat; reported by the same outlet (github.com).
NVIDIA supports Kubernetes / Shared entity: GPU / Earlier coverage / Tension
Linked by a graph relationship (NVIDIA supports Kubernetes); both cover GPU; earlier GPU coverage from 2026-02-22.
NVIDIA supports Kubernetes / Shared entity: Red Hat / What happened next
Linked by a graph relationship (NVIDIA supports Kubernetes); both cover Red Hat; picks up the Red Hat thread on 2026-08-05.
Linked by a graph relationship (NVIDIA supports Kubernetes); both cover Red Hat; picks up the Red Hat thread on 2026-07-27.