Reproducibility Audit of LLM-Driven Vulnerability Research: Only 55.6% of Artifacts Run, and Embedded Oracles Fire on Patched Builds 20 Times Out of 30
A pre-registered audit posted 2026-08-10 screened a 104-paper corpus (2023–2026) of LLM/agent-driven vulnerability validation work; only 59 papers (56.7%) have a publicly reachable artifact. Executing an 18-paper sample plus all 102 cases of an anchor benchmark, 58/102 (56.9%) anchor cases carry a script-internal CVE ID that diverges from the declared directory CVE, and only 10/18 artifacts complete their declared workflow at first run (11/18 after environment-only repair). Worst: artifact-embedded oracles show sensitivity 60% and specificity 45% — 20/30 patched-counterfactual audits still produce the claimed signal on the patched build. A trigger on a vulnerable build is not evidence of CVE-specific reproduction.
↳ Follow the thread