Skills
FIRE: harness-injected runtime policies at pre-failure states lift Terminal-Bench pass^2 from 64.4% to 73.6% without touching the model
FIRE mines states that preceded observed agent failures and has the harness inject targeted instructions or action denials when those states recur. GPT-5.6 Sol's best-of-two barely moved (+1.2 points) while repeated success rose 9.2 points, so the policies turn solutions the agent can already reach into ones it delivers every time. In a randomized five-arm test, real policies scored 61% against 36-43% for a timing-matched sham and generic 'verify' or 'reconsider' nudges. Generic self-check prompts do almost nothing. State-specific rules taken from your own failure logs do.
↳ Follow the thread