Research
Good Token Hunting: Principled Token Selection Cuts Visual Geometry Transformer Costs
Visual geometry transformers for multi-view 3D reconstruction process all image tokens equally, wasting compute on uninformative regions. This paper provides a systematic guide to token selection strategies — which tokens to keep, which to prune — with concrete benchmarks on reconstruction quality vs. compute trade-offs. Practitioners doing 3D reconstruction, novel view synthesis, or dense stereo matching can apply these selection strategies to reduce inference cost without quality loss.
Source
↳ Follow the thread