Fetching from the wire…
Public story · 2026-08-23 · high
Ten frontier models were tested on 202 attorney-annotated legal questions, and none topped an F2 score of 0.46 at catching what's missing.
Why now: Covered in the August 23 briefing, and none of the ten models tested have closed the gap yet.
InsufficiencyBench pushed ten frontier models through 202 attorney-annotated legal questions, and none scored above an F2 of 0.46 at catching what's missing, per the paper.
Median recall on missing facts landed at 0.44. That means the average model missed more than half of what practicing attorneys flagged as necessary before answering. That's a problem for anyone building an agent meant to answer legal questions instead of flagging what it doesn't know.
The benchmark spans six legal domains and 24 US jurisdictions. Attorneys wrote and annotated each question so it leaves out at least one legally material fact. The benchmark checks whether a model notices a missing fact and names it. It also checks whether the model withholds a conclusion until it has that fact.
The paper sorts the failures into two camps. Some models hedge on everything, qualifying every answer whether or not a fact is actually missing. Others just answer, filling the gap with an assumption the paper calls a fabricated presumption, and never tell the user they did it.
Bigger models won't close this gap on their own. I'd bet the models that hedge indiscriminately are the safer pick for legal use right now. That's true even though hedging tanks their benchmark score compared to models that just answer, confidently wrong. Worth watching whether any lab tries a dedicated pre-answer check instead of more scale.
The paper doesn't say whether a targeted prompt telling a model to check for gaps first would move these scores. It also only covers six legal domains, so nothing here says how models handle missing facts in medical or financial questions.
Each link below shares sources, entities, or timing with this story.
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (agent, annotated, frontier, model); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (agent, frontier, model); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (agent, frontier, model); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (agent, answer, model); pushes against this story (but).
Reported by the same outlet (arxiv.org); overlapping topics (agent, frontier, model); pushes against this story (against).
Same source domain / Shared topic / Downstream implication
Reported by the same outlet (arxiv.org); overlapping topics (agent, frontier, model); traces where this leads (implication).
Same source domain / Shared topic
Reported by the same outlet (arxiv.org); overlapping topics (agent, answer, conclusion, model).
Reported by the same outlet (arxiv.org); overlapping topics (answer, conclusion, fabricated, model).