Bayesian self-escalation lets an agent hand off mid-reasoning, beating post-hoc routing at equal cost
This 25 August paper (arXiv 2608.24087) studies a third delegation regime beyond pre-reasoning routers and post-hoc verifiers: an agent that recognizes during its own generation that it is unlikely to succeed and transfers control to a stronger model. Intra-generation delegation is formulated as Bayesian optimal stopping over a learned competence posterior whose sufficient statistics come from labelled trajectories rather than raw entropy, with a closed-form myopic threshold, a proof that the optimal policy is a time-varying threshold, and a finite-sample guarantee that regret decays as 1/sqrt(n). Real-model validation on a Qwen2.5-Coder 1.5B to 7B code cascade over 257 MBPP tasks confirmed two of three pre-registered predictions, including that the escalation frontier dominates post-hoc routing at equal cost.
Source
↳ Follow the thread