A Decades-Old Sparse-Approximation Algorithm Matches Every Purpose-Built Long-Video Frame Selector
Long-video models keep only a small fixed slice of frames (an hour at 1 fps is 3,600 images), and which frames survive is usually treated as preprocessing detail, with published selectors changing scorer, prompt boundary, resolution policy and answering model all at once. Holding each fixed and varying one decision at a time across six training-free rules, three benchmarks and two answering models, selection is the largest single lever: eight query-selected frames beat sixteen uniformly spaced ones by 6.9 points on LongVideoBench's hour-long bin, and unmodified Orthogonal Matching Pursuit matches or comes within a point of every purpose-built selector on all three benchmarks. Halving each frame's spatial budget at fixed timestamps costs at most 0.44 points, so compression is close to free and the savings can be reinvested.
↳ Follow the thread