Research
PACE Compresses Pixels Before the Vision Encoder, Keeping 93.8% of Qwen2.5-VL-7B on 10% of Visual Tokens
Most visual token pruning runs after the vision encoder, leaving the encoder's own latency untouched, and struggles to keep both global context and fine detail under tight budgets. PACE is training-free and works in two stages: an Adaptive Pixel Compressor scores visual information density before encoding and downsamples redundant input, then a Dynamic Dual-Attention Extractor keeps tokens using both internal encoder signals and semantic signals from the LLM. Integrated into Qwen2.5-VL-7B it retains 93.8% of original performance on 10% of visual tokens for a 3.1x speedup in time to first token, with code at github.com/jjL357/PACE.
↳ Follow the thread