Elastic Attention Cores for Scalable Vision Transformers
arXiv·medium signal
Alan Z. Song et al. propose Elastic Attention Cores that reduce the computational cost of all-to-all self-attention in Vision Transformers while preserving their data-driven scaling properties. The method enables ViTs to dynamically adjust attention computation based on input complexity, achieving strong accuracy-efficiency tradeoffs. Directly applicable for practitioners deploying vision models under latency or compute constraints.