VisPCO: Pareto-Optimal Visual Token Pruning Makes VLM Inference Faster Without Quality Guessing
arXiv·medium signal
VisPCO frames visual token pruning as a Pareto configuration optimization problem, automatically identifying optimal pruning configurations for vision-language models rather than relying on manually-tuned thresholds. The framework discovers computation-performance optimal points that existing fixed-pruning approaches consistently miss, offering practitioners a principled way to speed up VLM inference.