随着Agent AI的兴起和本地化小型模型部署的增加,CPU在大语言模型推理中的角色正在重新定义。传统训练工作负载中CPU与GPU的比例为1:8,但在推理阶段已提升至1:4,在Agent工作负载中更是趋向1:1,部分客户的部署比例甚至达到4:1。[1]
业界领导者纷纷调整战略以适应这一转变。Intel首席执行官Lip-Bu Tan在Computex 2026表示:"对于强化学习、编排和Agent,CPU是更好的选择。"[1] 根据Intel Q1 2026财报,其Data Center and AI部门营收达51亿美元,同比增长22%。[1] Arm首席执行官Rene Haas估计,传统AI数据中心需要约3000万CPU核心/GW,而在Agent时代这一需求上升至1.2亿CPU核心/GW,增长了4倍。[1]
CPU在Agent工作负载中的重要性源于其在工具处理上的优势。Georgia Tech与Intel的联合研究表明,Agent工作负载中CPU侧工具处理占总端到端延迟的50-90%。[1] 硬件层面,NVIDIA推出的Vera CPU相比主流x86 CPU性能大幅提升,沙箱性能快1.8倍,内存带宽快2倍,单核带宽快3倍。[1] NVIDIA的NVL72机架架构体现了这一转变,配置72个Rubin GPU和36个Vera CPU,CPU:GPU比例从传统的1:8转变为1:2。[1]
市场增长前景广阔。OpenAI与AWS在2025年11月签订了为期7年、价值380亿美元的基础设施合作协议,涵盖数百万NVIDIA GPU和数千万CPU。[1] 摩根士丹利预计,Agent CPU的这一转变将在2030年前产生325亿至600亿美元的增量CPU市场增长。[1]
The landscape of AI infrastructure is shifting as companies recognize the expanded role of CPUs in large language model deployment. According to Intel's Q1 2026 financial results, the CPU-to-GPU ratio has evolved significantly across different workload types, moving from 1:8 in training environments to 1:4 for inference tasks, and approaching 1:1 or even 4:1 in agent-based applications [1]. This rebalancing reflects the growing complexity of AI systems that rely on tasks beyond pure matrix multiplication, where CPUs demonstrate distinct advantages.
The emergence of Agent AI as a dominant workload category has catalyzed hardware innovations across the industry. Research conducted by Georgia Tech and Intel reveals that CPU-side tool processing accounts for 50–90 percent of total end-to-end latency in agent workloads [1], underscoring the computational significance of these tasks. Intel CEO Lip-Bu Tan stated at Computex 2026 that "for reinforcement learning, orchestration, and agents, the CPU is the better choice" [1]. Meanwhile, NVIDIA has introduced the Vera CPU, which delivers 1.8x faster sandbox performance, 2x greater memory bandwidth, and 3x higher single-core bandwidth compared to mainstream x86 processors [1]. The company's NVL72 rack architecture exemplifies this architectural shift, pairing 72 Rubin GPUs with 36 Vera CPUs for a 1:2 CPU-to-GPU ratio, a dramatic departure from the traditional 1:8 configuration [1].
Market projections suggest substantial growth ahead. Arm CEO Rene Haas estimated that traditional AI data centers require approximately 30 million CPU cores per gigawatt, while the agent era will demand 120 million CPU cores per gigawatt—a fourfold increase [1]. Morgan Stanley projects that this CPU transition will generate $32.5–60 billion in incremental CPU market growth before 2030 [1]. The scale of this infrastructure buildout is evident in major strategic commitments: OpenAI and AWS signed a seven-year infrastructure partnership worth $38 billion in November 2025, encompassing millions of NVIDIA GPUs and tens of millions of CPUs [1]. Intel's Data Center and AI division reported Q1 2026 revenue of $5.1 billion, representing 22 percent year-over-year growth [1].