Every AI-Generated Automation Script Tested Contained Exploitable Vulnerabilities — Across ChatGPT, Copilot, and Gemini Alike
Shanna M. Kahn and John D. Hastings collected code from ChatGPT, Microsoft Copilot, and Google Gemini using identical prompts across three automation domains, then had Claude Code perform a standardized vulnerability review scored with CVSS v3.1 and mapped to OWASP Top 10:2021 and MITRE ATT&CK. Every script contained exploitable vulnerabilities; nine of the 17 identified vulnerability classes appeared in output from all three models, 14 of 17 in at least two, and weighted CVSS scores differed by under 10% across platforms. Their conclusion is the useful one for teams shopping vendors: risk tracks the task category, not the model, so the question is whether LLM-generated automation ships without review at all.
↳ Follow the thread