Fetching from the wire…
Public story · 2026-07-25 · high
The gains come from a 3:1 mix of linear and softmax attention, which the authors say is the reusable part, not the video demo.
Why now: It's part of the July 25 research briefing, where the 120x throughput comparison against Wan 2.2-A14B is drawing as much attention as the VBench score.
SANA-Video 2.0 generates a 720p, five-second clip in 13.06 seconds on a single GPU, according to the paper posted to arXiv. That's the fully optimized number; the compiled DiT forward pass alone runs 3.2 times faster than a full-softmax build at the same resolution.
That speed matters because attention is the expensive part of video diffusion transformers, and the cost climbs fast with sequence length, exactly what 720p frames produce. SANA-Video 2.0 handles most of that mixing with gated linear attention, an O(N) approach, and drops in gated-softmax attention as an anchor every fourth block, a 3:1 ratio. The authors add what they call Block Attention Residuals on top, which raise the model's effective rank by about 12%.
Both the 5B and 14B versions fit on a single GPU. At 480p the model scores 84.30 on VBench, and the paper claims roughly 120 times the throughput of Wan 2.2-A14B, a full-softmax model in the same weight class.
The paper's own framing treats the video results as secondary to the method. The 3:1 linear-to-softmax ratio and the rank fix are architecture choices any diffusion transformer team can lift, video generation or not. Watch for other labs to start reporting their own linear-to-softmax ratios instead of just parameter counts.
Each link below shares sources, entities, or timing with this story.
In autoregressive video diffusion the growing KV cache makes attention the dominant inference cost, and existing eviction heuristics cause inter-frame flicker. HeadCast does a one-time classification at the maximum-noise step sorting every attention head into Sink, Dummy, Spat...
Seed round led by General Catalyst with Box Group, Emergence, Gradient and SV Angel, announced August 26. Founded by CEO Phillip Li, it recreates enterprise software including Salesforce, Workday and email clients as full digital twins with permission systems and webhooks inta...
Built after Indeed laid off the author's wife at seven months pregnant, the product has 4,300+ authed users, 91 paying, three people hired in four weeks, and ingests roughly 15,000 job listings daily straight from employer career pages, then classifies, enriches and embeds the...
TechCrunch's follow-up details the deal mechanics: led by ICONIQ with Sequoia reinvesting and a16z on the cap table, unplanned and inbound. Rillet rebuilt the general ledger rather than layering automation on an existing ERP, and has about 600 customers, most migrating off Ora...
Load Hijack modifies nothing but router weights in a checkpoint. When a private trigger appears, token-to-expert assignment concentrates on experts co-located on a single GPU, making it a straggler while peers idle (arXiv 2608.10614). Across three MoE families and four corpora...
June, founded by former Salesforce executive Efrat Rapoport with three co-founders from Bonobo AI (acquired by Salesforce in 2019), took money from Michael Dell, Aaron Levie and George Kurtz alongside Time Ventures. The product scans existing Salesforce, ServiceNow, Databricks...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.