Sources
PackForcing: 2-Minute Coherent Video on Single H200 GPU via Three-Partition KV Cache — 24x Temporal Extrapolation
PackForcing solves long-video generation with a bounded 4GB KV cache using a three-partition strategy: sink tokens (full-resolution anchor frames), mid tokens (32x spatiotemporal compression via 3D convolutions + VAE re-encoding), and recent tokens (full resolution for local coherence). Generates 2-minute 832x480 video at 16 FPS on single H200. Achieves 24x temporal extrapolation (5s training → 120s output) with state-of-the-art temporal consistency (26.07) and dynamic degree (56.25) on VBench.
↳ Follow the thread