研究人员对AMD最新GPU架构的矩阵乘法核心进行了系统性研究,开发了准确的数值模型1。该研究覆盖AMD CDNA 1、CDNA 2、CDNA 3三代架构,涉及MI100、MI210/250和MI300A/300X等GPU型号1。研究团队通过设计测试向量对这些GPU进行了特征刻画,开发了MATLAB软件模型,并利用1000万组随机测试向量验证了比特级可重现性1。
研究发现AMD矩阵核心在数值处理上存在多个特性差异,包括累加器宽度、舍入行为、规范化点、中间上溢/下溢逻辑、次规范数处理和特殊输入处理等方面1。研究结果表明AMD矩阵核心不符合IEEE 754浮点标准,且与NVIDIA张量核心存在应用级精度差异1。论文于2026年9月13日提交初稿(v1),随后于9月15日发布修订版本(v2)1。
Researchers have developed comprehensive numerical models characterizing the behavior of matrix multiplication cores in AMD's latest GPU architectures 1. The study systematically examined the CDNA 1, CDNA 2, and CDNA 3 generations across multiple GPU models, including the MI100, MI210/250, and MI300A/300X 1. By designing targeted test vectors and validating against 10 million random input sets, the team achieved bit-level reproducibility in their MATLAB software models 1.
The investigation revealed significant differences in how AMD's matrix cores handle numerical operations 1. The research identified variations in accumulator width, rounding behavior, normalization points, intermediate overflow and underflow logic, subnormal number handling, and special input processing across the different architectures 1. The findings demonstrate that AMD matrix cores do not conform to IEEE 754 floating-point standards and exhibit application-level precision differences when compared to NVIDIA tensor cores 1. The work was initially submitted on September 13, 2026, with a revision released on September 15, 2026 1.
评论
还没有评论,欢迎留下第一条。