Poolside发布了Laguna S 2.1模型,这是一个118B总参数的混合专家模型,其中8B参数在推理时处于激活状态,支持100万token的上下文窗口 1。该模型从训练开始到发布仅耗时不到9周 1。
在编码基准测试中,Laguna S 2.1表现突出。在Terminal-Bench 2.1上的得分达到70.2%(启用思维模式),相比未启用思维模式的60.4%有明显提升 1;在DeepSWE v1.1上的得分为40.4%,较基础版本的16.5%大幅改进 1。这些成绩展现了该模型在编码任务上的竞争力,超越了多个体型更大的模型。
模型的训练数据基础广泛,共包含409k个代理和非代理环境,其中83k用于终端任务,168k用于软件工程工作流 1。软件工程任务中最大的数据来源是约38,000个真实提交任务,分布在17,000个代码仓库中 1。
除了编码能力外,Laguna S 2.1还展现了独立推理的能力。该模型独立重新推导了Erdős问题#397,并发现了无穷解族(在知识截断日期2025年11月之后)1。在自我优化方面,模型成功改进了自身代理框架的性能,速度提升了5.2%,内存分配降低了70% 1。
在OpenRouter平台上,该模型的定价为输入token $0.10、输出token $0.20,缓存读取每100万token $0.01 1。
Poolside has unveiled Laguna S 2.1, a 118-billion-parameter mixture-of-experts language model designed to excel at extended reasoning and long-horizon tasks 1. The model features 8 billion active parameters and supports a context window of 1 million tokens 1.
The development cycle proved remarkably efficient, with the model advancing from initial training to release in under 9 weeks, beginning on May 22 1. Despite its scale, Laguna S 2.1 demonstrates competitive performance on coding benchmarks compared to significantly larger models 1.
Performance and Capabilities
The model achieved a 70.2% score on Terminal-Bench 2.1 when operating in reasoning mode, up from 60.4% in standard operation 1. On DeepSWE v1.1, it reached 40.4%, a substantial improvement from the baseline 16.5% 1. These gains highlight the effectiveness of the model's reasoning capabilities 1.
The training dataset encompassed 409,000 agent and non-agent environment interactions, with 83,000 examples focused on terminal tasks and 168,000 on software engineering workflows 1. Software engineering training drew from approximately 38,000 real commit tasks distributed across 17,000 repositories 1.
Noteworthy Achievements
Laguna S 2.1 independently re-derived Erdős problem #397 and discovered an infinite family of solutions, following a knowledge cutoff in November 2025 1. The model also successfully optimized its own agent framework, achieving a 5.2% speed improvement while reducing memory allocation by 70% 1.
On OpenRouter, pricing is set at $0.10 per million input tokens and $0.20 per million output tokens, with cached reads charged at $0.01 per million tokens 1.
评论
还没有评论,欢迎留下第一条。