Wafer AI在AMD MI355X显卡上成功部署了Kimi K3大模型。[1]Kimi K3拥有2.8万亿参数,模型权重需要超过1.5TB显存。[1]原本需要16张NVIDIA B200跨越两台服务器运行的模型,现在仅需8张MI355X在单台服务器上完成。[1]8张MI355X合计约2.3TB显存。[1]
在性能对标方面,MI355X展现出明显优势。[1]MI355X峰值总吞吐量达952 Token/s,而16张B200双节点部署总吞吐量仅为498 Token/s,单节点吞吐量约为B200双节点部署的3.8倍。[1]在成本效率上,MI355X每美元提供约48 Token/s的峰值吞吐量,相比B200约7 Token/s和B300约33 Token/s更具竞争力。[1]MI355X与B300均拥有288GB单卡显存。[1]
AMD为该模型部署提供了首发级的软件支持。[1]Wafer AI通过推测解码优化使单路性能提高2.2倍,并通过补零将12个注意力头补到16个,优化预填充速度从4000-7000 Token/s提升到1.3万Token/s。[1]
Wafer AI has successfully deployed the Kimi K3 large language model on AMD MI355X GPUs, demonstrating significant efficiency gains over NVIDIA's B200 processors [1]. The 28 trillion-parameter model, which originally required 16 NVIDIA B200 GPUs distributed across two servers, now runs on just 8 MI355X units within a single server [1]. With model weights exceeding 1.5TB in memory requirements, the 8 MI355X GPUs provide a combined 2.3TB of VRAM, sufficient to handle the deployment on one machine [1].
Performance metrics reveal MI355X's substantial advantage in throughput. The 8-GPU MI355X configuration delivers a peak aggregate throughput of 952 tokens per second, compared to 498 tokens per second for the 16 B200 GPU dual-node setup [1]. Additionally, MI355X achieves approximately 48 tokens per second per dollar in peak throughput efficiency, significantly outperforming B200 at 7 tokens per second per dollar and B300 at 33 tokens per second per dollar [1]. Both MI355X and B300 feature 288GB of memory per card [1]. AMD provided first-tier software support for the deployment, enabling the model to run with minimal optimization requirements [1].