Fetching from the wire…
Infra2026-08-11 · source-backed
arXiv 2608.08097 exploits decode-time attention sparsity: keep a 2,048-token budget resident and speculatively prefetch the rest via lookahead prediction. Reported results are 1.69x speedup on reasoning workloads, up to 2.1x on multi-GPU long-context serving, roughly 2x throughput under prefill-decode disaggregation, and 2.2–2.6x less decode-node host memory. Accuracy lands within 0.7 points of full attention overall. If this replicates, long-context serving cost is a memory-hierarchy problem, not a hardware-budget one.
Each link below shares sources, entities, or timing with this story.
H100 uses HBM / Shared entity: GPU / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (H100 uses HBM); both cover GPU; overlapping topics (accuracy, long-context).
Shared entity: Accuracy / Same source domain / Shared topic / Earlier coverage / Tension
Both cover Accuracy; reported by the same outlet (arxiv.org); overlapping topics (accuracy, cost).
H100 uses HBM / Shared entity: GPU / Shared topic / Earlier coverage
Linked by a graph relationship (H100 uses HBM); both cover GPU; overlapping topics (cache, cost).
Shared entity: GPU / Same source domain / Shared topic / Earlier coverage
Both cover GPU; reported by the same outlet (arxiv.org); overlapping topics (accuracy, attention, cache).
Both cover GPU; reported by the same outlet (arxiv.org); overlapping topics (attention, cache, cost).
Shared entity: HBM / Same source domain / Shared topic / Earlier coverage
Both cover HBM; reported by the same outlet (arxiv.org); overlapping topics (cache, cost, serving).
Shared entity: GPU / Same source domain / Shared topic / Earlier coverage
Both cover GPU; reported by the same outlet (arxiv.org); overlapping topics (accuracy, cache).
Shared entity: Accuracy / Same source domain / Shared topic / Earlier coverage
Both cover Accuracy; reported by the same outlet (arxiv.org); overlapping topics (accuracy, cost).