Mistral's Shieldstral: a 3B Policy-Adaptive Safety Classifier Trained on 54.1M Samples That Beats Models 7x Its Size
An eleven-author paper (arXiv 2607.25857, July 28) including Guillaume Lample and Giada Pistilli introduces Shieldstral, a 3B-parameter multimodal safety classifier that matches or beats models nearly 7x larger on text safety benchmarks and sets a new state of the art on multimodal safety classification. The design trick is reformulating all of content moderation as a single binary yes/no question-answering task, which lets heterogeneous safety datasets with incompatible taxonomies be consolidated under one training framework rather than requiring a taxonomy merge. The paper publishes the full data construction recipe covering curation and generation of roughly 54.1 million samples plus a fine-grained evaluation set specifically for measuring policy adaptability, which matters because most guardrail models cannot be retargeted to a new policy without retraining.
↳ Follow the thread