一项开源项目展示了在7个ESP32S3微控制器组成的集群上运行BitNet 1.58位量化语言模型的分布式推理能力1。该项目采用主从架构,由1个主节点和6个计算节点组成,通过高速SPI菊花链实现节点间通信1。
该系统运行的是0.5B规模的语言模型,采用BitNet三值量化方案,将嵌入层量化为INT4格式并存储在约14MB的Flash中1。主节点负责处理BPE分词和嵌入层计算,6个计算节点共同承载24层Transformer块,实现了极限压缩下的模型分布式推理1。项目提供了完整的软件栈,包括主板固件、节点固件和Python量化工具,并以MIT许可证开源发布1。
An open-source project has demonstrated distributed inference of a 0.5 billion parameter language model across a cluster of seven ESP32S3 microcontrollers, employing extreme quantization techniques to achieve operation on resource-constrained hardware.1 The system comprises one primary node and six computing nodes that coordinate through high-speed SPI daisy-chain communication.1
The architecture divides computational tasks across the cluster, with the main node handling byte-pair encoding tokenization and embedding layers, while the six computational nodes collectively execute the model's 24 Transformer blocks in four-layer configurations per node.1 The embedding layer is quantized to INT4 precision and requires approximately 14 megabytes of flash storage.1 Communication between nodes occurs via optimized SPI daisy-chain protocol to synchronize inference across the distributed system.1
The implementation utilizes BitNet's 1.58-bit ternary quantization scheme, enabling the model to operate within the severe memory constraints of microcontroller environments.1 The complete project includes firmware for both the primary and compute nodes, along with Python-based quantization tools, and has been released under the MIT License.1
评论
还没有评论,欢迎留下第一条。