Cactus团队发布了Needle 3自动化模型,这是一款参数量在25至121百万之间、经2-bit量化后仅8-29MB的超小型语言模型1。该模型采用了Monarch Hadamard MLP等创新架构设计,支持工具调用和结构化JSON输出功能1。在Mobile Actions基准测试中,20层模型达到86.0分的性能表现1。
Needle 3展现出了出色的边缘设备适配能力1。在Raspberry Pi 5上,该模型的解码速度最高可达4k tokens/sec,预填充速度达到10k tokens/sec1。模型支持英语、法语、西班牙语、德语、荷兰语、意大利语和波兰语等7种语言1,并兼容包括macOS、Linux(x86-64、ARM64、ARMv7、RISC-V、MIPS32)、Windows、Android、iOS、watchOS、tvOS、浏览器WebAssembly和WASI组件在内的多个平台部署1。
The Cactus team has unveiled Needle 3, a series of automation models ranging from 25 to 121 million parameters that achieve performance comparable to DeepSeek V4 Flash despite their minimal size 1. When quantized to 2-bit precision, these models compress to just 8-29 MB, enabling deployment on resource-constrained devices 1. The architecture incorporates innovative techniques including Monarch Hadamard MLP and supports tool calling and structured JSON output 1.
Benchmark results demonstrate the models' competitive capabilities across multiple scales 1. A 20-layer variant achieved 86.0 points on the Mobile Actions benchmark, outperforming comparisons including LFM2.5 1.2B at 82.4 points and Qwen3.5 0.8B at 76.0 points 1. On Raspberry Pi 5 hardware, the models deliver decoding speeds up to 4,000 tokens per second, with prefill speeds reaching 10,000 tokens per second 1.
The Needle 3 suite supports multiple languages including English, French, Spanish, German, Dutch, Italian, and Polish 1, and runs across a comprehensive platform range spanning macOS, Linux on multiple architectures (x86-64, ARM64, ARMv7, RISC-V, and MIPS32), Windows, Android, iOS, watchOS, tvOS, browser-based WebAssembly, and WASI components 1.
评论
还没有评论,欢迎留下第一条。