开源项目openTPU展示了人工智能系统自主开发硬件加速器的能力1。该项目完成了一个完整的AI推理加速器的设计和实现,包括硬件设计、指令集、编译器和主机软件等全套组件1。这一加速器已在Xilinx Kintex-7 xc7k480t FPGA卡上成功验证,配备两个DDR3通道1,能够运行包括LFM2-2.6B、SmolLM3-3B、Phi-4-mini、Qwen3.5等十个现代语言模型1。
在性能方面,该硬件加速器实现了显著的优化。相比之前版本,最新构建的解码速度提升了8-10%,同时DRAM利用率从82-87%提升至91-94%,达到DDR3-1066峰值带宽17.1 GB/s的82-94%1。此外,加速器支持4-bit量化权重,相比int8方案的解码速度提升幅度达到40-45%1。该项目包含完整的开源工具链,涵盖SystemVerilog硬件设计、ISA模拟器、内核语言编译器和主机软件1。
The open-source project openTPU has demonstrated artificial intelligence's ability to autonomously design hardware accelerators, creating a complete AI inference system that encompasses hardware design, instruction set architecture, compiler, and host software.1 The accelerator has been validated on an FPGA card and is capable of running multiple modern language models while achieving DRAM utilization rates of 82 to 94 percent of DDR3 peak bandwidth.1
The openTPU accelerator operates on a Xilinx Kintex-7 xc7k480t FPGA card equipped with two DDR3 channels and supports the execution of ten contemporary models, including LFM2-2.6B, SmolLM3-3B, Phi-4-mini, and Qwen3.5.1 A recent build iteration improved decoding speed by 8 to 10 percent compared to the previous version while increasing DRAM utilization from 82-87 percent to 91-94 percent.1 The system leverages 4-bit quantized weights, which deliver a 40 to 45 percent speed increase in decoding performance relative to int8 approaches.1 The complete implementation includes an open-source toolchain comprising SystemVerilog-based hardware design, an ISA simulator, kernel language compiler, and host software.1
评论
还没有评论,欢迎留下第一条。