Stealthy Backdoor Attacks Against LLMs Using Natural Style Triggers
arXiv·medium signal
New backdoor attack method against LLMs using natural writing style as trigger patterns instead of explicit tokens, enabling attacks that preserve text naturalness while reliably injecting attacker-specified payloads in long-form generation. Prior methods suffered from unnatural trigger patterns, unreliable payload injection, and incompletely specified threat models. This attack is harder to detect because the triggers are stylistic rather than lexical.