德国公司Aleph Alpha于2026年10月3日推出了开源大语言模型Kolibri1。这是一个为德语和英语优化的专家混合模型,采用Apache 2.0许可证1。模型拥有78亿总参数,但每个token仅激活约3.5亿参数1,支持最多1,048,576 tokens的上下文窗口1。该模型在德国和芬兰基础设施上从零开始训练1,符合欧盟AI法案要求1。
在性能表现上,Kolibri在多项德语任务中领先同规模开源模型。在德语AIME 2025数学测试中,该模型得分为87.5分1。在德语知识保留测试中,Kolibri的得分达到70.8,高于Qwen3.5 35B的69.81。在德语法律文本处理中,Kolibri的tokenizer效率比GPT-5减少约15%的token数1。模型在Omniscience测试中承认不知道的比率达到44%,而Qwen3.5为11.1%1。GPU内存方面,该模型需要约78GB容量,可运行在两张NVIDIA A100/H100或单张H200/B200/B300上1。
Aleph Alpha, a German company, released Kolibri, an open-source large language model optimized for German and English, on October 3, 2026 1. The model features a mixture-of-experts architecture with 7.8 billion total parameters but activates approximately 350 million parameters per token, reducing computational overhead while maintaining performance 1. Kolibri was trained from scratch on infrastructure based in Germany and Finland, aligning with European Union AI Act requirements 1.
The model demonstrates competitive capabilities within its size class. It supports a context window of up to 1,048,576 tokens and achieves 87.5 points on the German AIME 2025 mathematics test, the highest score among open-source models of comparable scale 1. Additionally, Kolibri's German-language tokenizer proves more efficient than GPT-5, requiring 15 percent fewer tokens when processing legal documents 1. On knowledge retention tests specific to German content, the model scores 70.8, outperforming Qwen 3.5's 35 billion parameter variant at 69.8 1. The model demonstrates epistemic caution, acknowledging unknown information 44 percent of the time in the Omniscience test, compared to 11.1 percent for Qwen 3.5 1. Kolibri is distributed under the Apache 2.0 license and requires approximately 78 gigabytes of GPU memory, compatible with two NVIDIA A100 or H100 GPUs or a single H200, B200, or B300 1.
评论
还没有评论,欢迎留下第一条。