Skills
Agents hit 80% F1 deciding whether a dependency CVE is exploitable, but fall below 70% explaining why
VEX-Bench is the first benchmark for the actual supply-chain security question, whether a known upstream vulnerability is reachable in a downstream project, as distinct from prior zero-day discovery benchmarks. It contains 75 real-world cases mined from GitHub and labeled by security experts across Python, Java and Go, evaluated over nine models and three agent harnesses. GPT-5.5 and Claude Opus 4.6 reach roughly 80% F1 on binary vulnerability-status classification, but only GPT-5.5 clears 70% macro-F1 on fine-grained justification, meaning an agent can tell you a Dependabot alert is a false positive far more reliably than it can tell you the correct reason to record in a VEX document.
↳ Follow the thread