EmDash插件注册表已引入Cloudflare的Clef决策模型进行内容审核1。Clef是一个多模态AI模型,可同时处理文本和图像,通过回答关于文本/链接的9个问题和关于图标/截图的8个问题来检测钓鱼、冒充、骚扰等不当内容1。该系统的检查范围涵盖明确的性内容、仇恨/去人性化内容、图形暴力、钓鱼/凭证请求、冒充、诈骗、垃圾邮件、欺骗链接、误导性媒体/声明以及审核操纵尝试1。
新审核系统在性能和可靠性上表现突出1。当模型判定概率达到0.45或更高时,相关内容将转交人工审核1。在63次公开文本测试中,Clef在所有运行中都做出了预期的通过或审核决定,无无效输出或不一致1。端到端文本审核延迟为1.64秒(p95)1,同时该模型通过了所有37个插件配置文件和43个评估时注册表中的实时列表图像1。相比之下,前一代系统需要两个通用模型在文本内容上达成一致,另需单独进行图像处理,而新系统用一个多模态模型替代了这一复杂流程,审核速度更快且实现更简洁1。
EmDash has implemented Cloudflare's Clef, a multimodal AI model, to automatically moderate its plugin registry 1. Clef processes both text and images simultaneously, evaluating submissions through nine questions about text and links alongside eight questions about icons and screenshots 1. The system identifies problematic content including explicit sexual material, hate speech, graphic violence, phishing attempts, impersonation, fraud, spam, deceptive links, misleading media, and moderation manipulation 1. Flags with a confidence score of 0.45 or higher are escalated to human reviewers for final determination 1.
This deployment streamlines the moderation workflow compared to the previous approach, which required multiple models to reach consensus on textual content before undergoing separate image analysis 1. In testing, Clef demonstrated consistent performance across 63 public text samples, producing the expected pass or review decisions in all runs with no invalid outputs or inconsistencies 1. The system achieved an end-to-end text moderation latency of 1.64 seconds at the 95th percentile 1. Additionally, Clef successfully processed all 37 plugin profile images and 43 live registry list images during evaluation 1.
评论
还没有评论,欢迎留下第一条。