CatchBench Scores 72 Auditors Across 1,187 Configs and Publishes 71 of 118 Contrasts as Unresolved Rather Than Ranked
CatchBench asks when an agent failure can be caught given three information states: the declared configuration before a run, a growing prefix of its trace, and the finished trace. It scores 72 entrants, from rule scanners to eleven LLM judges across nine model families, over 1,187 declared configurations and 1,162 recorded runs under seven task contracts with their own labels rather than one leaderboard. Only 47 of 118 pre-declared contrasts separate, and the sharpest result cuts against the authors' own data: a rule that ignores every name and permission and just flags each capability declared after the first reaches perfect F1 on one of six configuration sources, measuring how the corpus was built rather than how well a method reasons.
↳ Follow the thread