尽管DeepSeek通过算法优化降低了算力依赖,但随着大模型能力提升和应用场景扩展,算力需求反而激增,形成"杰文斯悖论"[1]。超节点技术——将多张AI芯片融为一体——成为解决这一困境的最优方案[1]。2025年国内智能算力规模为1590 ExaFLOPS,预计到2026年上半年达2185 EFLOPS,同比增长177%[1]。
国产AI芯片厂商已纷纷推出超节点产品,在大规模推理场景中实现高效率、低成本运营[1]。百度昆仑芯片收入预计从2025年约13亿元飙升至2026年的83亿元,同比增长超6倍[1]。平头哥真武AI芯片累计出货56万片,在智驾行业已部署超13万卡[1]。昇腾推出的384卡超节点已销售750多套,目前已迭代至1024卡超节点,并将推出8192卡超节点[1]。
从技术角度看,互联带宽成为关键瓶颈。CHIP实验室主任罗国昭指出:"如果互联带宽不足,大量GPU不得不停下来等待数据同步,即使理论算力充足,也难以真正发挥出来"[1]。沐曦股份研发副总裁黄向军表示:"超节点这个架构比普通的服务器更适合万亿参数大模型"[1]。
Despite algorithmic optimizations that have reduced computational demands, the rapid expansion of large language model capabilities and application scenarios has paradoxically intensified pressure on computing infrastructure, creating what analysts describe as a "Jevons Paradox" in artificial intelligence.[1] Domestic AI chipmakers have identified supernode technology—which integrates multiple AI chips into a unified system—as the optimal solution to address this capacity shortage.[1]
The Chinese computing market is experiencing accelerated growth, with intelligent compute capacity projected to reach 1,590 ExaFLOPS in 2025 and expand to 2,185 ExaFLOPS by mid-2026, representing a year-over-year increase of 177 percent.[1] Baidu's Kunlun chip revenue is anticipated to surge from approximately 1.3 billion yuan in 2025 to 8.3 billion yuan in 2026, exceeding a sixfold increase.[1] Meanwhile, Ping Tou Ge's Zhenwu AI chip has achieved cumulative shipments of 560,000 units, with more than 130,000 cards already deployed in the autonomous driving sector.[1] Ascend's supernode architecture has sold over 750 units with 384 cards each and has already advanced to a 1,024-card iteration, with plans to introduce an 8,192-card supernode.[1]
The efficiency gains from supernode architecture are particularly evident in large-scale inference operations. Kimi K3, featuring parameters totaling 2.8 trillion, activates only 104 billion parameters per individual token.[1] According to Muxi Technology's Vice President of Research and Development Huang Xianjun, "this supernode architecture is better suited to trillion-parameter large models compared to conventional server configurations."[1] CHIP Laboratory Director Luo Guozhao emphasized the critical importance of interconnect bandwidth, noting that "when interconnection bandwidth is insufficient, numerous GPUs must pause and wait for data synchronization; even with theoretically adequate computational power, it becomes difficult to fully realize that capacity."[1]