研究人员提出了持久状态机(PSM)框架,用于在FPGA上实现大语言模型的注意力操作[1]。该框架采用INT4量化的内存单元设计,通过在Zynq-7000和UltraScale+两个硬件平台上的实现验证,展示了低功耗和高集成度的特性[1]。
在性能指标方面,该架构实现了极低的功耗表现[1]。动态功耗控制在1.0毫瓦以下,归一化动态能耗为3.81 × 10^-5皮焦/操作[1]。在Zynq-7000平台上完成了1024单元阵列(d=128)的实现,而在UltraScale+ xcvu9p平台上的256单元子阵列以62.5兆赫系统时钟成功完成时序闭合,建立时间余量为+1.854纳秒[1]。片上系统仅占据设备逻辑切片的0.67%和数字信号处理块的0.00%[1]。通过对超过一千个随机向量的功能仿真验证,该设计与定点软件参考实现达到了位精度一致[1]。该研究已递交日本专利申请(No. 2026-177318)[1]。
Researchers have introduced the Persistent State Machine (PSM) framework, designed to implement attention operations for large language models on FPGA hardware using INT4-quantized memory cells [1]. The architecture has been validated on two platforms—the Zynq-7000 and UltraScale+—demonstrating significant efficiency gains in power consumption and resource utilization [1].
The implementation achieves dynamic power consumption below 1.0 mW with normalized dynamic energy consumption of 3.81 × 10⁻⁵ pJ/op [1]. On the Zynq-7000 platform, a 1,024-cell array with dimension d=128 was successfully deployed [1], while a 256-cell subarray on the UltraScale+ xcvu9p achieved timing closure at a 62.5 MHz system clock with a worst negative slack of +1.854 ns [1]. The system-on-chip footprint remains minimal, occupying only 0.67% of device logic slices and 0.00% of DSP blocks [1].
Functional validation through simulation of over one thousand random vectors confirmed bit-accurate alignment with fixed-point software reference implementations [1]. A patent application has been filed in Japan under application number 2026-177318 [1].