Grayscale and Color Inversion Bypass All Three Major Commercial Image-Moderation APIs
'Old Tricks, New Models' (arXiv 2607.28187, July 30) runs a large-scale black-box evaluation of three commercial foundation-model-based image moderation services against seven simple, model-agnostic transformations requiring no gradients, surrogate models, or knowledge of the target. All three can be bypassed, and even fixed transforms like color inversion and grayscale conversion flip unsafe-to-safe decisions while leaving content plainly recognizable to humans. Robustness varies sharply by dataset and harm category, with multimodal content and self-harm most vulnerable — the practical conclusion being that swapping task-specific classifiers for a foundation-model API is not itself a security boundary and must sit inside a layered pipeline.
Source
↳ Follow the thread