HoF-Bench: A Deliberately Minimal LLM Analyzer Rediscovers 68% of Real AI-Found CVEs Without Any Frontier Model
HoF-Bench (arXiv 2607.27030, 2026-07-29) is built from 95 public CVEs that AISLE's analyzer discovered across eight repositories, pinned at their vulnerable commits, with a detector-blinded frontier judge that credits a finding only if it matches code path, root cause, attack condition, and impact. A minimal scaffold rediscovers up to 65 of 95 (68%) using only five open-weight models (21B–284B total, 3–13B active) and five proprietary small/flash-tier models — no frontier model performs detection anywhere in the study, across 7,600 model–CVE pass records. For builders this reframes the cost model of automated vulnerability hunting: the scaffold, not the model tier, is doing the work, and every CVE missed by all ten models concentrates in C infrastructure code.
↳ Follow the thread