Research
On Success and Simplicity: A Second Look at the Transferable Vision-Language Attack Pipeline
Recent transferable attacks on vision-language pretraining models pile on complicated loss functions and multi-stage image/text perturbations, but this paper shows a simpler pipeline is both cleaner and more successful. The authors identify three overlooked issues from inappropriate cross-modal interactions and excessive operations, then remove them to boost transfer success. A cautionary, useful read for teams red-teaming or hardening multimodal models against adversarial transfer.
Source
↳ Follow the thread