在芯片短缺的背景下,AI行业正将焦点转向充分挖掘现有GPU的潜力。[1]卡内基梅隆大学研究发现,GPU整体上有近20%的执行时间和约11%的能耗被浪费在等待上,推理场景问题更为突出——Azure Code负载有65%能耗消耗在空转上,OpenAI Chat类请求达52%。[1]一台NVIDIA GB200 NVL72机柜采购成本约400万美元,若软件栈优化不足致GPU利用率仅50%,相当于浪费200万美元。[1]
围绕这一问题,开源推理引擎SGLang和vLLM相继推向商业化,融资规模均超1亿美元。[1]孵化SGLang的RadixArk公司推出了包括KV Cache复用、Prefill/Decode分离等优化技术方案。[1]Anthropic通过KV Cache优化直接将成本砍了90%,其推理队伍拥有200多人,在三年内从基础设施极度不稳定发展到稳定,并实现了二季度盈利。[1]
市场对GPU利用率优化的需求推动了相关企业的快速增长。[1]Baseten一年收入增长20倍,估值从21亿美元暴涨至130亿;fireworks七个月估值翻四倍达175亿美元,年化营收突破10亿美元。[1]这种增长背景是全球AI基础设施投资的持续扩大——Meta、Google、Microsoft等科技巨头今年合计资本开支预计超1万亿美元,黄仁勋预测到2030年,全球AI基础设施年投资规模将达4万亿美元。[1]
The AI industry is confronting significant chip shortages, prompting major technology companies and startups to focus on optimizing the utilization of existing GPUs.[1] Research from Carnegie Mellon University has revealed that GPUs waste approximately 20% of execution time and around 11% of energy on idle waiting, with inference workloads experiencing even more severe inefficiencies—Azure Code loads waste 65% of energy on idle operations, while OpenAI Chat requests lose 52%.[1]
The economic stakes of underutilized hardware are substantial. A single NVIDIA GB200 NVL72 cabinet costs approximately $4 million to procure, and when software optimization is insufficient, GPU utilization rates hover around 50%, representing a $2 million loss per cabinet.[1] Open-source inference engines SGLang and vLLM have emerged as leading solutions to this challenge, with both securing seed-round funding exceeding $100 million.[1] RadixArk, the company that incubated SGLang, has deployed technical solutions including KV Cache reuse and Prefill/Decode separation, alongside a reinforcement learning training framework called Miles, with the goal of raising GPU utilization from current levels of approximately 50% to over 90%.[1]
Companies implementing these optimizations are experiencing remarkable results. Anthropic achieved a 90% cost reduction through KV Cache optimization alone, while maintaining a 200-plus person inference team that transformed from severe infrastructure instability three years ago to achieving profitability in the second quarter.[1] The commercial sector shows explosive growth, with Baseten experiencing 20-fold annual revenue growth and its valuation soaring from $2.1 billion to $13 billion, and fireworks seeing its valuation quadruple to $17.5 billion within seven months while surpassing $1 billion in annualized revenue.[1] Meanwhile, major technology firms including Meta, Google, and Microsoft are collectively projected to invest over $1 trillion in capital expenditures this year, with industry projections suggesting global AI infrastructure investments will reach $4 trillion annually by 2030.[1]