InsufficiencyBench: no frontier model exceeds F2 = 0.46 at spotting what a query is missing
arXiv 2608.20220 (20 Aug 2026) tests whether models recognize that a question omits legally material facts, name what is missing, and withhold a conclusion. It defines eight missing-element categories across three failure modes (switch, gating, fatal prerequisite) and builds 202 items (58 base queries, 144 deficient variants) over six legal domains and 24 US jurisdictions, annotated by practising attorneys. Across ten frontier models none exceeds F2 = 0.46 on missing-element identification, median recall is 0.44, and models either hedge indiscriminately or answer silently on fabricated presumptions. The generalizable point for agent builders is that clarification-before-action is not an emergent behavior you can assume.
Source
↳ Follow the thread