Agents
The best Anthropic model catches 47% of expert-identified requirement defects while false-flagging 11%
arXiv 2609.03230 (2026-09-03) benchmarks ten off-the-shelf models across two families and five generations each, over 100 independent runs, two requirement sets and five sampling temperatures, against expert ground truth built on INCOSE criteria. The error profile is strongly asymmetric — a 47% median detection rate at 11% false flags for the best performer — and necessity and correctness issues, the ones needing engineering judgment, are almost always missed. Generational progress was non-monotonic, so newer model versions cannot be assumed better at this task.
Source
↳ Follow the thread