Agents
The handoff tax: escalating a stuck coding agent to a stronger model recovers less than half the quality gap and costs a premium
Posted 25 August (arXiv 2608.24358), this study switches models mid-run on long coding tasks using low-cost/high-cost pairs from the Claude and GPT families, varying direction, timing, and how much trajectory the receiver inherits. Full-trajectory escalation from the cheap model to the strong one recovers under half of the quality gap while incurring a substantial cost premium, which the authors name the handoff tax; downshifting after the hard reasoning is done is the favorable direction. The preferred interface reverses with direction: giving the strong model less of the weak model's trajectory improves escalation, while removing the strong model's trajectory hurts downshift.
Source
↳ Follow the thread