Skills
Long-horizon agents have an identifiable 'submit early' pressure state you can read off hidden states
A linear probe detects when a tool-use agent is biased toward submitting a polished-looking final answer with key constraints unresolved, and activation interventions along that direction change whether the agent keeps calling tools or stops. Controlled context manipulation shows the pressure is reduced by constraint clarity and explicit action mapping. The practical version for anyone without hidden-state access is the ablation result: stating constraints explicitly and mapping them to actions in the prompt measurably delays premature finalization.
↳ Follow the thread