AI编程Agent公司Infinity近日用其开发的Ignition工具为芯片公司d-Matrix仅花10小时搭建了类CUDA软件,这一事件再次引发业内对英伟达CUDA护城河的讨论[1]。同时,DeepSeek也开源了TileKernels GPU算子库,通过TileLang编程语言来减少对手写CUDA代码的依赖[1]。
CUDA由英伟达高管Ian Buck带队花近20年打造,涵盖内核层、库层、框架层三个层次[1]。然而业内观点认为,推理市场正成为撬动CUDA的缺口[1]。Rebellions首席商务官Marshall Choy表示,"推理侧CUDA不再是决定因素,竞争转向开源软件"[1]。INT21创始人Bing Xu也指出,虽然"Agent能在短时间内生成大量代码,但验证才是最大瓶颈"[1]。对此,英伟达开发者生态副总裁Ankit Patel回应称,"我们也在使用AI Agent来更快地开发CUDA"[1]。
Infinity当前估值已达1亿美元,已融资1500万美元[1]。成立于2019年的d-Matrix主要专注生成式AI推理芯片的开发[1]。
Artificial intelligence is rapidly eroding the competitive advantages that NVIDIA has spent nearly two decades building around its CUDA software ecosystem. AI programming agent company Infinity recently demonstrated this vulnerability by using its Ignition tool to construct CUDA-like software for chip manufacturer d-Matrix in just 10 hours [1]. This achievement has sparked industry debate about the durability of NVIDIA's once-dominant position in GPU software infrastructure.
The challenge to CUDA extends beyond Infinity's speed demonstration. DeepSeek has released an open-source TileKernels GPU operator library written in TileLang, offering an alternative path for developers seeking to reduce reliance on hand-written CUDA code [1]. Industry observers note that inference applications have become the critical pressure point. According to Marshall Choy, chief commercial officer at Rebellions, inference no longer depends on CUDA as a decisive factor, with competition shifting toward open-source software solutions [1]. CUDA was built over nearly 20 years under the leadership of NVIDIA executive Ian Buck and encompasses three architectural layers: kernel-level, library-level, and framework-level components [1].
However, NVIDIA's entrenched position remains difficult to displace completely. Bing Xu, founder of INT21, acknowledges that while AI agents can generate vast amounts of code rapidly, verification represents the true bottleneck [1]. NVIDIA itself is adapting to this landscape; Ankit Patel, vice president of developer ecosystems at NVIDIA, stated that the company is also employing AI agents to accelerate CUDA development [1]. Infinity, which has raised $15 million in funding and achieved a current valuation of $100 million, continues to push this frontier [1], while d-Matrix, founded in 2019 and focused on generative AI inference chips, represents the type of emerging competitor seeking to bypass traditional CUDA dependencies [1].