Fermion Research 推出了 Neutrino-1 8B,一款参数量为 8.19B 的开源大语言模型 [1]。该模型采用专有的三进制权重格式,将文件大小压缩至 3.88 GB [1],可在数据中心 GPU、Mac 和桌面 CPU 上运行 [1]。
模型在不同硬件上的推理性能表现各异 [1]。在 Apple M5 CPU 上,单流推理速度达到 24.9 tokens/秒(9 线程配置);NVIDIA L4 GPU 可实现 30.7 tokens/秒;Apple Silicon MacBook 的推理速度最快,达到 33.7 tokens/秒 [1]。在投机式解码任务中,0.6B 草稿模型在计数提示上的接受率达到 100%,在事实提示上的接受率为 96.5% [1]。
该模型衍生自 Qwen3-8B [1],已在 Apache License 2.0 下开源发布,允许商业使用、修改、微调和再分发 [1]。
Fermion Research has unveiled Neutrino-1 8B, an open-source large language model with 8.19 billion parameters, featuring a proprietary ternary weight format that compresses the model file size to 3.88 GB [1]. The compact model enables deployment across diverse hardware platforms, from data center GPUs to Mac systems and desktop CPUs [1].
The model demonstrates competitive inference performance across multiple configurations [1]. On Apple M5 CPUs with nine threads, Neutrino-1 8B achieves 24.9 tokens per second, while NVIDIA L4 GPUs deliver 30.7 tokens per second with a 4.68 GB context window [1]. Apple Silicon MacBooks reach 33.7 tokens per second [1]. In speculative decoding benchmarks, a 0.6B draft model shows a 100% acceptance rate on counting prompts and 96.5% on factual prompts [1].
The model is released under the Apache License 2.0, permitting commercial use, modification, fine-tuning, and redistribution [1]. Neutrino-1 8B is derived from Qwen3-8B, which similarly uses the Apache-2.0 license [1].