一家AI公司发布了Mercury 2.5大语言模型,在性能上实现了显著突破。1该模型相比前代产品Mercury 2在智能水平上提升40%,同时推理速度达到1107个令牌每秒,并支持260K令牌的上下文窗口。1
定价方面,该模型的启动价为输入令牌百万个0.04美元、输出令牌百万个0.15美元,后续标准价格调整为输入令牌百万个0.20美元、输出令牌百万个0.75美元。1
Mercury 2.5已在生产环境中广泛应用于搜索、语音和编码产品。1公司同时推出了Mercury Voice和Mercury Router两款预览产品,其中Mercury Voice的首个令牌响应时间低于170毫秒。1
根据实际应用案例,使用该模型的OpenCall平台实现了显著的性能改进,P99响应时间从数分钟降低至1秒,P50响应时间从0.4秒降至0.2秒以下。1编码产品Augment Code的性能优化更为突出,压缩延迟从约150秒降至27秒,改进幅度达82%,同时成本下降90%。1
下一代模型正在训练中,预计将在未来数月内发布。1
Inception has unveiled Mercury 2.5, a large language model that delivers a 40% improvement in intelligence compared to its predecessor Mercury 2 1. The new model achieves a processing speed of 1,107 tokens per second on NVIDIA GPUs and supports a context window of 260,000 tokens 1.
The pricing structure for Mercury 2.5 starts at $0.04 per million input tokens and $0.15 per million output tokens during launch pricing, with standard rates set at $0.20 per million input tokens and $0.75 per million output tokens 1. The model is already deployed in production environments across the company's search, voice, and coding products 1.
Real-world performance gains demonstrate the model's practical impact. In the OpenCall case, P99 response times decreased from several minutes to one second, while P50 response times improved from 0.4 seconds to below 0.2 seconds 1. For Augment Code, compression latency was reduced from approximately 150 seconds to 27 seconds—an 82% improvement—while costs dropped by 90% 1. Mercury Voice, one of the preview products released alongside the model, delivers its first token in under 170 milliseconds 1.
Inception has signaled that the next generation of Mercury models is currently in training, with plans to release them within the coming months 1.
评论
还没有评论,欢迎留下第一条。