Fetching from the wire…
Public story · 2026-02-22 · source-backed
Tsinghua University's hybrid Top-k + Top-p masking with distillation fine-tuning achieves 95% attention sparsity and 16.2x attention speedup on video diffusion models while maintaining quality. This is what makes real-time AI video generation practical — direct cost and latency improvements for production systems. arXiv
Each link below shares sources, entities, or timing with this story.
Shared entities / Same source / Shared topic / What happened next
Both cover SpargeAttention2, Sparsity, Speedup; cite the same source (arXiv); overlapping topics (generation, sparsity).
Shared entity: SpargeAttention2 / Same source / Shared topic / Earlier coverage
Both cover SpargeAttention2; cite the same source (arXiv); overlapping topics (attention, cost, diffusion, generation, video).
Shared entity: Speedup / Same source domain / Shared topic / What happened next
Both cover Speedup; reported by the same outlet (arxiv.org); overlapping topics (achiev, cost, speedup).
Shared entity: Sparsity / Same source domain / Shared topic / What happened next
Both cover Sparsity; reported by the same outlet (arxiv.org); overlapping topics (cost, maintaining, sparsity).
Shared entity: SpargeAttention2 / Same source / What happened next
Both cover SpargeAttention2; cite the same source (arXiv); picks up the SpargeAttention2 thread on 2026-02-23.
Shared entity: SpargeAttention2 / Same source / Earlier coverage
Both cover SpargeAttention2; cite the same source (arXiv); earlier SpargeAttention2 coverage from 2026-02-20.
Shared entity: SpargeAttention2 / Same source
Both cover SpargeAttention2; cite the same source (arXiv).
Same source domain / Shared topic
Reported by the same outlet (arxiv.org); overlapping topics (attention, cost, diffusion, video).