Fetching from the wire…
Public story · 2026-07-30 · high
The bugs it missed all cluster in C infrastructure code, not spread evenly across the eight repositories.
Why now: As of July 30, no frontier model has been run through HoF-Bench, leaving the C-code gap untested.
A minimal scaffold rediscovered 65 of 95 real-world CVEs, or 68%, per arXiv 2607.27030. Those 65 are the ones AISLE's analyzer first found across eight repositories. The scaffold matched them using models running as few as 3 billion active parameters.
HoF-Bench pins each CVE at its vulnerable commit and grades results with a detector-blinded judge. The judge only credits a match on code path, root cause, attack condition, and impact together. Ten models ran the scaffold: five open-weight, sized 21B to 284B total parameters, and five proprietary models at small or flash tiers.
Across 7,600 model-CVE pass records, the paper's authors say the scaffold explains the results, not the model tier.
Every CVE that none of the ten models found sits in C infrastructure code. The paper doesn't say why C resists rediscovery more than the rest of the corpus. It also doesn't test whether a frontier model would close that gap.
Each link below shares sources, entities, or timing with this story.
CWEAgent, built on a structured representation capturing root cause, trigger condition, violated property, exploit mechanism and impact, scored 85% top-1 on a curated 100-CVE benchmark, then audited 15,556 open-source CVEs disclosed 2017 through 2026 (arXiv 2608.21977). Just u...
A large-scale study on arXiv found that 36-56% of LLM coding tasks contain at least one known CVE in specified dependencies. Not in the generated code itself. In the packages the model tells you to install. The numbers get worse. 62-75% of those CVEs are rated Critical or High...
The Model Context Protocol has a security problem that's no longer theoretical — it's statistical. Between January and February 2026, researchers filed 30+ CVEs against MCP servers, clients, and infrastructure. One package with nearly 500,000 downloads carried a CVSS 9.6 RCE....
Deng et al. built 120 real-case-grounded tasks across 20 business scenes in six financial domains, running four self-evolving scaffolds on a shared Qwen3.7-Max backbone against paired non-evolving controls. Letta posted the highest evolved score (91.65) and fewest compliance i...
The agent skills supply chain is under coordinated attack. Snyk's ToxicSkills audit found 36% of ClawHub's 3,984 skills contain prompt injection payloads, 13.4% have critical malware, and submission rates exploded 10x to 500+/day. This week alone: CVE-2026-2256 (CVSS 9.1) is a...
Kai Security mapped all 30 CVEs into three attack layers: execution (43% — exec()/shell injection), tooling (20% — infrastructure attacks), and new attack classes (14% — eval() injection, env var injection). The flagship CVE is CVE-2026-0755 (Gemini MCP Tool, CVSS 9.8) with pu...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.