Skills
Code-generation guardrails fail almost completely on code-to-code requests, and a fictional-scenario wrapper defeats them at near 100%
CS-Guard evaluated 9 guardrails across seven LLMs using 1,000 malware-generation prompts, 7 jailbreak attacks and 331 code-to-code prompts spanning infilling, completion and translation. After jailbreaks, average attack success rate for text-to-code reached about 50% for many guardrails; for code-to-code it approached 100% on base models and stayed between 14.4% and near 100% across guardrails. A new fictional scenario attack that embeds malicious intent inside a legitimate software-development story achieved close to 100% ASR across many guardrails. If you rely on a guardrail layer in a code pipeline, the code-to-code path is the one it almost certainly does not cover.
↳ Follow the thread