Skills
An automated red-teamer plants instruction backdoors in customized coding LLMs at 94.5% success while evading detection entirely
ARIA generates instruction backdoor attacks against customized coding LLMs automatically, reaching a 94.5% attack success rate while preserving clean-task utility and evading detection at false negative rates reported up to 100%. The threat surface is the customization layer itself — system instructions and configuration you paste into a GPT-style wrapper or shared assistant, not the base model weights. If your team shares custom coding assistants, the instruction blob needs the same code review as the code it writes.
↳ Follow the thread