Compute Where it Counts: Self-Optimizing Language Models Allocate Per-Token Compute Dynamically
arXiv·medium signal
SOL pairs a frozen LLM with a lightweight policy network that reads hidden states and selects per-token efficiency actions — jointly controlling attention sparsity, structured MLP pruning, and activation quantization bit-width. Unlike static compression, SOL adapts compute budget to token difficulty without modifying base model weights, offering a practical path to 'pay only for hard tokens.'