Skills
Stop reasoning when the action stabilizes: a convergence probe plus perplexity deferral cut thinking compute 43–73%
TSDS pairs a convergence probe that halts local reasoning once the intended action stops changing with a perplexity-based rule that escalates only genuinely uncertain actions to a larger cloud model, jointly tuned by Learn-Then-Test to give finite-sample guarantees on both expected episode reward and cloud-call rate. It cut per-episode thinking compute 43–73% on three of four benchmarks (HotpotQA, MBPP, a household robot task) versus deferral-only baselines. The generalizable technique for anyone running a small-model-first tier: gate escalation on action stability rather than a fixed reasoning budget, and calibrate the two thresholds together instead of separately.
↳ Follow the thread