Automated Framework to Harden LLM System Instructions Against Encoding Attacks
arXiv·medium signal
Addresses system instruction leakage — a critical risk in OWASP Top 10 for LLM Applications where instructions may contain API credentials, internal policies, and privileged workflow definitions. Proposes automated evaluation and hardening framework specifically targeting encoding-based attacks that bypass instruction protection without requiring expensive reasoning models. Practical for anyone deploying agents with sensitive system prompts.