ResearchZeroDayBench frontier LLMs vs zero-daysWweb·medium signalXBlueskyLinkedInCopy linkGPT-5.2, Claude Sonnet 4.5, Grok 4.1 tested against 22 novel critical vulnerabilities. Frontier LLMs cannot autonomously discover/patch zero-days. ICLR 2026 Agents in the Wild workshop. arXiv 2603.02297↳ Follow the threadPolicy dependency / Stack layerCROSS-CATEGORY: Three Independent Agent-Action Gates Shipped in 48 Hours, All Judging the Command Against Stated IntentProduct Hunt, github.com/AGGIB/Stroq and rewarelabs.com (three independent sources; the 72% figure is Reware's own)Policy dependency / Stack layerPattern: the agent config layer is being treated as an unmanaged dependency graph, and three independent sources said so this weekarXivStack layer / Threat pattern16% of 3,171 public agent-harness setups carry a confirmed security defect, and 3.8% ship a skill that pre-approves your shellarXivPolicy dependency / Stack layerNSA, CISA and FBI Name Six Chinese AI Firms in a Joint Advisory on Industrial-Scale DistillationCISAStack layer / Threat patternAgentAudit attaches to a running agent and scores its trace on ten dimensions, exposing 95.1 vs 22.6 trust spreads at similar task completionarXivStack layer / Threat patternVibe Coding Cut Task Time 27% and Raised Security Vulnerabilities in the Same TrialarXiv 2609.09560Policy dependency / ContrastA Fine-Tuned 4B Qwen in 2.6 GB Beats GPT-5.6 on a Transit-Kiosk Agent Benchmark, and PEFT Gains Vanish by 27BarXiv 2609.10016Stack layer / ContrastMicrosoft's September patch is the largest on record at roughly 972 CVEs, with two zero-days already exploitedArs Technica (counts corroborated by BleepingComputer, CrowdStrike and Cybersecurity News)