Multimodal Jailbreaks Aren't Missing Safety — They Shift Representations Outside an Existing Refusal Boundary
MMAligner (arXiv 2608.05909, Aug 6) explains why multimodal LLMs refuse an unsafe text prompt yet answer the semantically equivalent image-plus-text version. A geometric analysis of representations finds that safety learned from text does persist across modalities — a shared safety subspace and refusal boundary remain effective — but unsafe multimodal inputs undergo a representation shift that lands them outside the boundary. Calibrating those representations back into the pre-existing refusal region with a hard lower bound, soft upper bound, and benign-input preservation objective raises average refusal rate on unsafe multimodal inputs to 99% with under 2% utility degradation and minimal training data.
↳ Follow the thread