Research
Answer Accuracy Hides Invalid Reasoning: Correct-Answer/Invalid-Trace Rates Stay 45.8-59.1% on BIRD Mini-Dev
Dutta and Moharir argue that a benchmark-correct answer from a structured-data agent can rest on a computation that never actually happened, and define Trace Integrity as a deployment criterion requiring the recorded computation be explicit, executable, schema-valid, operator-faithful, replayable, answer-consistent, and auditable. On BIRD Mini-Dev, Direct SQL, Operation Summary + SQL, and Contract-First SQL reach answer accuracies of 20%, 22%, and 24% while their Trace Integrity pass rates are 39%, 43%, and 40%, and their CAIT (correct answer, invalid trace) rates stay at 55%, 59.1%, and 45.8%. Answer accuracy, trace validity, and silent-failure risk turn out to be three separate signals.
↳ Follow the thread