Research
HeadCast: Sort Attention Heads Into Four Archetypes Once, Get 1.95× Faster 1080P Video Generation With No Retraining
In autoregressive video diffusion the growing KV cache makes attention the dominant inference cost, and existing cache-eviction heuristics cause inter-frame flicker. HeadCast is training-free and plug-and-play: after a short warm-up it performs a one-time classification at the maximum-noise step sorting every head into Sink, Dummy, Spatial, or Global, then routes each through a head-specific cache pathway — critically retaining Global heads that carry long-range temporal consistency. Because the Spatial pathway runs on a fixed-size grid, savings scale with resolution: 1.62× at 720P and 1.95× at 1080P with VBench quality on par with full attention. Code at github.com/sjlgaga/HeadCast.
↳ Follow the thread