Fireworks Research推出了Ember-1模型,这是基于Kimi K3优化开发的专门版本1。该模型在保持相同质量的前提下,能将token消耗减少40%1。Kimi K3的推理token占生成token的90%以上1,Ember-1针对这一特点进行了专门优化。该模型通过50多次训练实验和200多次评估开发而成1,学会了高效推理能力,特别适用于编码和多轮对话任务1。
在实验评估中,Ember-1表现突出1。该模型在七个基准测试中的推理token缩短35-50%,同时未出现精度损失1,并在Doximity的Bedside Bench基准上创造了新的Pareto前沿,超越了GPT-5.6 Sol、GPT-6 Astra和Claude Opus 51。两个客户的生产环境A/B测试也验证了其效能,结果显示token节省约35%,而任务完成率和成功率保持不变或有所提高1。其中一名客户已在生产环境中运行Ember-1,并计划将其扩展为基础模型替代品1。
Fireworks Research has released Ember-1, a specialized variant of Kimi K3 engineered to reduce token consumption by 40% while maintaining equivalent output quality 1. The model represents the culmination of over 50 training experiments and more than 200 evaluation iterations designed to optimize inference efficiency 1.
In production testing, Ember-1 demonstrated substantial practical gains 1. Two customers running A/B tests in their production environments observed approximately 35% token savings, with task completion rates and success metrics remaining stable or improving 1. One customer has already deployed Ember-1 in production and plans to scale it as a replacement for their foundational model 1. Across seven benchmark evaluations and the two customer deployments, inference tokens were reduced by 35–50% without sacrificing accuracy 1.
The model achieved a new Pareto frontier on Doximity's Bedside Bench benchmark, surpassing GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5 1. Ember-1 is particularly effective for coding tasks and multi-turn conversations 1. This optimization addresses a critical inefficiency in Kimi K3, where inference tokens account for over 90% of total generated tokens 1.
评论
还没有评论,欢迎留下第一条。