Mage-VL Cuts Visual Tokens Over 75% by Encoding Motion Vectors and Residuals Instead of Sampling Frames
Microsoft's Mage-VL (arXiv 2607.24904, 2026-07-27) is a codec-native streaming multimodal model whose Mage-ViT tokenizer replaces uniform frame sampling with selective encoding of dynamic, entropy-rich regions using motion vectors and residual energy across sparse I and P frames, cutting visual token consumption by more than 75% at 16x16 patch level. Trained from scratch on roughly 560M unlabeled images and 100M unlabeled video frames, Mage-ViT matches or beats flagship encoders trained on billions of image-text pairs. Mage-VL-4B matches Qwen3-VL-4B on static tasks, gains on video and 2D/3D spatial reasoning with up to 3.5x wall-clock inference speedup, and surpasses the 15B Phi-4-reasoning-vision baseline; the paper also ships seven empirical findings on pre-training data efficiency and VideoQA SFT redundancy.
↳ Follow the thread