Research
ReBalance: Training-Free Efficient Reasoning for Large Reasoning Models — Accepted ICLR 2026
ReBalance is a training-free framework addressing both overthinking and underthinking in Large Reasoning Models using confidence metrics and steering vectors derived from hidden-state prototypes to dynamically guide reasoning depth. Tested across 9 benchmarks and 4 model sizes from 0.5B to 32B parameters, it reduces output redundancy while improving accuracy with no additional training cost. Directly applicable to practitioners using o1-style reasoning models for cost reduction.
Source
↳ Follow the thread