Agents
Outcome-only judges catch 84% of loud agent faults and 45% of silent ones, and no judge reads the final reply
Using a deterministic tool-using support-desk environment with a scripted oracle policy and a fault injector that breaks exactly one thing at a known step, the authors score five judges over 400 trajectories stratified by whether the customer-visible outcome survived the fault. The production-default outcome-only judge catches 84% of loud faults but only 45% of silent ones while flagging 33% of correct trajectories; a step-rubric judge reaches 77% silent recall with zero false alarms at 3x the cost. An invented promise appended to an otherwise perfect trajectory evaded the rule-based judge entirely and the step judge 82% of the time, and self-consistency tripled cost while improving nothing.
Source
↳ Follow the thread