PrismML推出了Ternary Bonsai 2 27B模型,该模型基于阿里巴巴的开源Qwen 27B模型12,采用三值量化技术实现显著的模型压缩。经过压缩后,模型大小仅为5.9GB,相比全精度版本缩小9至10倍12,使高性能推理能够在个人电脑和智能手机等本地设备上运行。
在性能表现方面,压缩后的模型在基准测试中达到原始Qwen模型98%的聚合基准得分2,相较2024年3月发布的第一代Bonsai模型的95%有所提升2。该模型支持262K-token的上下文窗口1和多模态输入1。在具体推理速度上,Bonsai 2 27B在NVIDIA RTX 5090上达到143 tokens/秒,在Apple M5 Max上达到46.8 tokens/秒1。能效方面,该模型在RTX 4090上的能耗为0.714 mWh/token,相比全精度8B模型的能效提高40%1。
技术上,模型采用三进制权重(包含+1、-1、0三个值)和FP16分组缩放,有效位宽达到1.76 bits/weight1。该模型已在NVIDIA GPU(CUDA)和Apple设备(MLX)上获得支持,并基于Apache 2.0许可证发布1。
PrismML此前已获得2225万美元种子轮融资,投资者包括Khosla Ventures、Cerberus Capital和加州理工学院2。第一代Bonsai模型已被下载超过1100万次,其中较小版本另有260万次下载2。公司CEO Babak Hassibi表示,团队计划在未来几个月内发布几百亿参数规模的模型2。
PrismML has unveiled Ternary Bonsai 2 27B, a compressed language model that reduces storage requirements to 5.9GB while maintaining 98.2% of the original performance capabilities 12. Built on Alibaba's open-source Qwen 27B model, the new version achieves a 9 to 10-fold reduction in memory footprint compared to the uncompressed original 12. This represents a significant improvement over the first-generation Bonsai model, which retained 95% of performance when released in March 2.
The compression is accomplished through ternary weight quantization, a technique that simplifies each weight to one of three values—positive one, negative one, or zero—combined with FP16 group scaling, resulting in an effective bit width of 1.76 bits per weight 1. The model supports a 262K-token context window and multimodal inputs, enabling efficient inference on consumer hardware 1. Testing on an NVIDIA RTX 5090 demonstrates throughput of 143 tokens per second, while Apple's M5 Max achieves 46.8 tokens per second 1. On an RTX 4090, the model consumes 0.714 milliwatt-hours per token, delivering 40% better energy efficiency compared to a full-precision 8B model 1. The model runs on both NVIDIA GPUs via CUDA and Apple devices via MLX, and is distributed under the Apache 2.0 license 1.
PrismML has secured $22.25 million in seed funding from investors including Khosla Ventures, Cerberus Capital, and the California Institute of Technology 2. The company's first-generation Bonsai model has already been downloaded over 11 million times, with smaller versions attracting an additional 2.6 million downloads 2. CEO Babak Hassibi indicated plans to release models with tens of billions of parameters within the coming months 2.
评论
还没有评论,欢迎留下第一条。