Distill your thinking-mode wins into short advisory prompts so the cheap non-thinking run inherits them — and they transfer across models
Rather than paying for reasoning mode on every call, this approach first identifies cases where thinking mode in a low-resource model avoided an error, has a larger model examine the error and its avoidance to produce explanations, then has a large model compress those explanations into brief advisory prompts — nuggets like "before coding, restate the requirements to clarify them." It is a mechanized version of the Socratic loop a student goes through: identify the mistake, reflect on the lapse, infer a rule, internalize it. The method works across many modest-sized models and the learned advisories sometimes transfer usefully to other models, and the paper also characterizes which classes of coding error the technique helps with — a concrete recipe for turning your own agent's expensive successes into a cheap standing prompt.
↳ Follow the thread