英国人工智能安全研究所(UK AISI)与美国AI标准创新中心(CAISI)近日联合发布了对月之暗面公司Kimi K3模型网络能力的初步评估报告。1
评估结果
在漏洞利用基准测试(ExploitBench)中,Kimi K3获得32%的得分,超过了同类开源模型GLM-5.2的24%。1该测试涵盖了41个2023年后发现的V8引擎漏洞。1
但在高级网络攻击模拟测试"The Last Ones"(TLO)中,Kimi K3的表现明显低于顶尖美国模型。1根据测试数据,Kimi K3在该攻击链中平均完成至第17步(总共32步),而美国顶尖模型平均完成28.5步。1在10次尝试中,Kimi K3仅成功完成1次。1
在任意代码执行(ACE)能力方面,Kimi K3在41个样本测试中达成0次,而最顶尖模型平均达成20次。1评估结果表明Kimi K3未能在任何任务中实现任意代码执行。1
产品发布计划
Kimi K3已于2026年7月16日发布,计划于2026年7月27日发布开源版本。1
The UK Artificial Intelligence Safety Institute (UK AISI) and the US Center for AI Safety and Innovation (CAISI) have jointly evaluated the cyber capabilities of Kimi K3, a model developed by Moonshot AI. 1
The assessment measured performance across multiple cybersecurity benchmarks. On the ExploitBench vulnerability exploitation benchmark, Kimi K3 achieved a score of 32%, outperforming the open-source GLM-5.2 model, which scored 24%. 1 However, in advanced attack simulation testing known as "The Last Ones" (TLO), Kimi K3 performed below leading American models. 1
Notably, Kimi K3 failed to achieve arbitrary code execution (ACE) in any test, registering zero successful instances across 41 samples, while leading American models averaged 20 successful ACE instances. 1 In the TLO attack chain assessment, Kimi K3 completed an average of step 17 out of 32 total steps, compared to an average of step 28.5 for top American models. 1 The model achieved only one successful completion in 10 attempted runs of the TLO scenario. 1
The ExploitBench testing covered 41 V8 engine vulnerabilities discovered after 2023. 1 Kimi K3 was released on July 16, 2026, with an open-source version scheduled for release on July 27, 2026. 1
评论
还没有评论,欢迎留下第一条。