ResearchDefensive Refusal Bias Safety Alignment Fails Cyber Defenders 2.72xarXiv·high signalXBlueskyLinkedInCopy linkSafety-tuned LLMs refuse legitimate defensive security tasks at 2.72x rate, system hardening 43.8% refusalSourceSource pagearXiv↳ Follow the threadStack layer / Threat pattern16% of 3,171 public agent-harness setups carry a confirmed security defect, and 3.8% ship a skill that pre-approves your shellarXivPolicy dependency / Stack layerMemSentry gates persistent memory writes on a signed security-state delta rather than on content classificationarXivStack layer / Threat patternPrivEscalate Scales Linux Privilege-Escalation Evaluation From Under 15 Scenarios to 531 Dockerized OnesarXiv 2609.09087Stack layer / Threat patternVibe Coding Cut Task Time 27% and Raised Security Vulnerabilities in the Same TrialarXiv 2609.09560Stack layer / Threat patternModels Spot Only 9.6% of Implementation Gaps in Research Specs but Fix 80.6% Once You Point Them OutarXiv 2609.10539Stack layer / Threat patternCode-generation guardrails fail almost completely on code-to-code requests, and a fictional-scenario wrapper defeats them at near 100%arXiv 2609.09798Stack layer / Threat patternAgentAudit attaches to a running agent and scores its trace on ten dimensions, exposing 95.1 vs 22.6 trust spreads at similar task completionarXivStack layer / Threat patternLLMs Claim to Find Bugs in Bug-Free Programs, and a Steering Vector Controls the Urge to EditarXiv 2609.10123