The Best LLM Catches 47% of Expert-Identified Requirement Defects While False-Flagging 11%
This is the first benchmark of off-the-shelf LLMs on requirements quality assessment against an expert-derived ground truth built on INCOSE criteria, covering ten models across two families (OpenAI and Anthropic) and five generations each, one hundred independent runs, two requirement sets and five sampling temperatures. The error profile is strongly asymmetric: the best-performing Anthropic model detects a median of only 47% of expert-identified issues while false-flagging 11% of clean requirements. Performance degrades further exactly where systems-engineering judgment is required, with necessity and correctness issues almost always missed, which puts a hard ceiling on treating an LLM as a review-cycle replacement rather than a first-pass filter.
↳ Follow the thread