A survey formalizes AI slop in vulnerability assessment and proposes measuring the gap with a Deductive Coverage Score
arXiv 2608.25667 (26 Aug) surveys the empirical evidence on hallucinated vulnerabilities, plausible but incorrect patches and semantically repackaged bug reports, arguing the load these impose on human triage functions like a denial-of-service on the pipeline. It traces the root cause to the gap between the causal deductive reasoning of security experts and autoregressive probabilistic generation, operationalizes that gap as a Deductive Coverage Score, and shows chain-of-thought prompting and tool-using agents narrow but do not close it. It also argues detection and watermarking are the wrong target because they establish provenance rather than correctness, and specifies two evaluation instruments, CVE-Bench and Slop-Score.
Source
↳ Follow the thread