Fetching from the wire…
Research2026-06-07 · source-backed
Multi-turn serving hurts because the KV cache grows linearly with conversation length, choking GPU memory and bandwidth. Tangram spends memory unevenly across the cache instead of treating all tokens equally, cutting the footprint of long sessions. Source: arXiv Directly useful if you self-host a model behind a chat or agent loop and your sessions run long.
Each link below shares sources, entities, or timing with this story.
Shared entity: GPU / Same source domain / Shared topic / What happened next
Both cover GPU; reported by the same outlet (arxiv.org); overlapping topics (agent, cache, cutting, memory).
Shared entity: Multi / Same source domain / Shared topic / What happened next / Tension
Both cover Multi; reported by the same outlet (arxiv.org); overlapping topics (agent, memory).
Shared entity: Directly / Same source domain / Shared topic / Earlier coverage / Tension
Both cover Directly; reported by the same outlet (arxiv.org); overlapping topics (agent, directly).
Both cover Directly; reported by the same outlet (arxiv.org); overlapping topics (agent, directly).
Shared entities / Same source domain / What happened next
Both cover GPU, Multi; reported by the same outlet (arxiv.org); picks up the GPU thread on 2026-07-25.
Both cover Directly, GPU; reported by the same outlet (arxiv.org); picks up the Directly thread on 2026-07-20.
Shared entity: Directly / Same source domain / Shared topic / Earlier coverage
Both cover Directly; reported by the same outlet (arxiv.org); overlapping topics (agent, directly, serving).
Both cover Directly; reported by the same outlet (arxiv.org); overlapping topics (agent, directly, memory).