Agents
OWASP Top 10 for LLM Applications barely agrees with 7,714 real incidents
Two OWASP working-group members compared the expert-consensus Top 10 against a corpus of 7,714 LLM-security incidents drawn from CVE, GHSA, OSV and AIAAIC, with 6,639 labeled against a 20-entry taxonomy. Agreement between expert ranking and incident-derived ranking was weak, at Cohen's κ ≈ 0.20 with a 90% interval crossing zero, and the 2026 candidate list resolves this by weighting expert consensus 75% and incident data 25%. Four frontier classifiers were tested for auto-labeling and none beat a 0.863 balanced-accuracy baseline. The paper is an exploratory analysis, not the official OWASP release.
Source
↳ Follow the thread