Mistral Shieldstral: 3B Open-Weights Multimodal Moderation
Mistral AI released Shieldstral, a 3B-parameter open-weights multimodal moderation model that outputs toxicity and safety scores across multiple axes, trained on ~600K human-judged examples covering harassment, hate spee…