Research
LoRAScan Detects Backdoored Adapter Triggers at Inference by Watching Just 5% of LoRA Insertion Sites, Rejecting 98.49% of Malicious Inputs
Untrusted LoRA adapters are a supply-chain threat — a backdoored adapter can emit malicious code, propaganda, or covert ads on a hidden trigger, and merging the adapter into the base model dilutes the signal enough to defeat adapter-agnostic defenses. LoRAScan's observation is that roughly 5% of LoRA insertion sites stay stable on clean inputs but spike sharply in down-projection activations when a trigger fires. It identifies those low-variance sites before deployment, monitors them at inference, and rejects ~98.49% of malicious inputs with a small clean-input error rate — without modifying adapter parameters.
↳ Follow the thread