When2Think Rewards a Model for Answering Easy Questions Directly Without Losing Accuracy on Hard Ones
When2Think (arXiv 2609.19671, submitted 17 Sep 2026) frames efficient reasoning as instance-adaptive compute allocation, targeting the pattern where large reasoning models overthink easy problems and underthink hard ones. Its Instance-level Difficulty-Aware Control is a reward-shaping mechanism using pre-computed reference statistics for accuracy and token usage to regulate reasoning depth, combined with verifier-based rewards and batch-wise standardized advantages for stable critic-free optimization with no learned reward model and no online reference-model queries. The claim is that it avoids the efficiency tax of uniform length penalties and rigid routing, which cut computation on easy instances by giving up accuracy on hard ones.
↳ Follow the thread