Mistral Ships Shieldstral: a 3B Apache-2.0 Safety Classifier That Takes Plain-Language Policies at Inference Time Instead of a Frozen Taxonomy
Mistral released Shieldstral on 2026-08-04 — a 3B open-weights multimodal safety classifier that reframes moderation as policy-adaptive question answering: you write the moderation rule in plain language at inference time and get back a calibrated safety score from a single token, with no retraining and one interface covering text and images. Mistral claims it matches open guard models up to 7x its size on text safety and sets a new state of the art on multimodal moderation, covers 12 languages, and runs on a single 16GB NVIDIA GPU. It ships under Apache 2.0 on Hugging Face as Mistral's inaugural contribution to the NVIDIA-led Open Secure AI Alliance (52 partners, launched July 28, notably without OpenAI, Google, or Anthropic). For builders, this is the first credible drop-in replacement for hardcoded-category guardrail models when your policy differs per jurisdiction or per product surface.
↳ Follow the thread