据彭博行业研究报告,DeepSeek 9月发布的V4.1 Flash模型在LiveBench基准测试中获得81.1分,与美国顶尖AI模型的性能差距已缩小至3%1。这一进展相比5月份的9%差距和年初的15%差距实现了大幅收窄1。
V4.1 Flash在LiveBench全球排名中位列第六,成为该榜单中排名最高的中国模型,仅低于Anthropic的最高分83.4分1。彭博高级分析师罗伯特·利亚表示,这意味着DeepSeek模型的表现已经"可以与Anthropic、OpenAI的领先AI系统相媲美"1。
报告指出,中国AI模型的快速进步得益于两方面因素:AI专业能力的不断深化,以及研究人员针对国产硬件优化模型的能力1。在LiveBench评选出的前15个模型中,仅有3个来自中国1。
DeepSeek's V4.1 Flash model, released in September, achieved a score of 81.1 on the LiveBench benchmark, reducing the performance gap between leading Chinese and American artificial intelligence systems to approximately 3 percent, according to a Bloomberg Industry Research report 1. This represents a significant convergence compared to the 9 percent differential observed in May and the 15 percent gap measured at the beginning of the year 1.
The V4.1 Flash model ranks sixth overall on LiveBench and stands as the highest-ranked Chinese model in the benchmark 1. It trails Anthropic's leading score of 83.4 points 1. Bloomberg senior analyst Robert Lya stated that DeepSeek's performance has now become "comparable to the leading AI systems from Anthropic and OpenAI" 1. Among the top 15 models identified by LiveBench, only three originate from China 1.
The rapid advancement reflects progress in deepening AI specialized capabilities and researchers' ability to optimize models for domestic hardware, the report noted 1.
评论
还没有评论,欢迎留下第一条。