ResearchSCAFFOLD-CEGIS: 43.7% of LLM Code Iteration Chains Have Security RegressionsarXiv·high signalXBlueskyLinkedInCopy linkIterative refinement paradox: code improves functionally but security degrades silently. Counterexample-guided synthesis to maintain security invariants.SourceSource pagearXiv↳ Follow the threadPolicy dependency / Stack layerBuild Integration, Not Fuzzing, Is What Kills LLM Dynamic Analysis on Real Autonomous-Vehicle StacksarXiv 2608.13450Stack layer / Threat patternLLM Repair Agents Write Patches 122% Larger Than Developers — and Minimality Prompts Don't Fix ItarXiv 2608.13292Stack layer / Threat patternAgentic patches are 122% larger than developer patches; RECAP is a post-generation refiner that shrinks them without losing fixesarXivStack layer / Threat patternFormal Specs Inferred From Tests Alone, Without White-Box Access to the ImplementationarXiv 2608.13240Stack layer / Threat patternStop asking the model for concise patches — refine them afterward and cut bloat from +242% to +4% over human patchesarXiv 2608.13292Stack layer / Threat patternCAPRI Adds a Machine-Readable Edit Contract to LLM Isabelle Proof Repair — 6 of 144 Accepted Proofs Had Touched Protected TextarXiv 2608.13459Stack layer / Threat patternTwo Copies of the Same Model Co-Fail on 90% of Missions — Redundancy Math for Agents Is BrokenarXiv 2608.12895Stack layer / Threat patternIterative LLM Infrastructure-as-Code Repair Silently Breaks Security Checks in 3.3% of Scenarios; Stop at Iteration 3arXiv 2608.13404