Mistral推出了Shieldstral,一款参数量仅为3B的开源多模态内容审核模型[1]。该模型采用策略自适应问答框架设计,性能表现超越体量大7倍的同类开源审核模型[1]。
Shieldstral可在单张16GB显存的GPU上运行[1],支持在推理阶段通过自然语言动态输入自定义审核策略,无需重新训练即可统一处理文本和图像的安全评估任务[1]。该模型返回校准的连续安全分数,以yes/no概率的形式呈现评估结果[1]。Mistral以Apache 2.0许可证将其开源发布[1]。
Mistral has unveiled Shieldstral, a 3-billion-parameter open-source model designed for multimodal content moderation [1]. The model employs a policy-adaptive question-answering framework and delivers performance comparable to or exceeding that of open-source moderation systems seven times its size [1]. Shieldstral operates on a single 16GB GPU, making it accessible for widespread deployment [1].
A distinctive feature of Shieldstral is its support for dynamic, natural-language policy inputs at inference time, eliminating the need for retraining when policies change [1]. The model returns calibrated continuous safety scores expressed as yes/no probabilities rather than fixed classifications [1]. It handles both text and image safety assessments through a unified framework, combining multimodal capabilities with policy flexibility [1]. Released under the Apache 2.0 license, Shieldstral is available as an open-weights model for community use and integration [1].