Agents
GenIaC-SecBench gives LLM-generated infrastructure code a size-matched human baseline for the first time
Vulnerability density in Infrastructure-as-Code is strongly inverse to artifact size (Spearman rho -0.55, p < 1e-77), so unmatched comparisons measure file size rather than security. Scanning 1,196 model artifacts from 12 configurations across four vendors and 634 human-authored templates with Checkov, Trivy and KICS, every model configuration lands within 3.21x to 3.87x human vulnerability density once matched on declared resource count. The gap is worst on trivial tasks (4.9x at one resource) and narrows to 1.4x at twenty, which inverts the usual intuition about where to review generated infra code.
Source
↳ Follow the thread