清程极智联合创始人师天麾认为,尽管AI Agent的应用形态不断演进,但Agent基础设施(Infra)建设仍处于早期阶段[1]。师天麾表示,"Agent Infra长期看一定会变得重要,但现在还没有到真正爆发的阶段"[1]。当前最紧迫的投入方向并非构造新的概念层,而是提高Token生产和交付效率、优化缓存管理、增强服务可靠性等底层推理基础设施问题[1]。
模型迭代的加速给企业长期投入带来了新的挑战[1]。缓存效率的下降尤为值得关注——缓存命中率从90%下降至80%,对应未命中比例从10%增加到20%,由于未命中部分需要完整的Prefill计算,实际成本增幅往往远超数字变化幅度[1]。在当前的定价机制下,缓存命中部分的输入成本可能仅为重新计算的约十分之一[1]。可靠性方面,若任务需连续调用模型20次且每次成功率为99%,完整链路一次成功的概率约为82%[1]。
为评估推理服务现状,清程极智搭建了AI Ping评测平台,目前覆盖30多家服务商和600多个模型服务[1]。师天麾还指出,2026年以来部分国产芯片的闲置情况已明显改善,部分新型号在较大规模集群中甚至出现供应紧张[1]。
The landscape of AI Agent development continues to shift quickly, yet foundational infrastructure challenges persist as the primary bottleneck for scaling these systems. At the WAIC 2026 InfoQ media livestream, Shi Tianwu, co-founder of Qingcheng Extreme Intelligence, emphasized that while Agent infrastructure will prove critical in the long term, the sector has not yet reached a genuine inflection point for explosive growth.[1] Instead of developing new conceptual frameworks, Shi argued that the most pressing investments should target lower-level reasoning infrastructure—specifically improving token production and delivery efficiency, cache management, and service reliability.[1]
The volatility of model iterations poses a significant risk to enterprise adoption, as companies must carefully evaluate which capabilities should be embedded as permanent features in their Agent workflows.[1] Shi highlighted concrete efficiency challenges: when cache hit rates declined from 90 percent to 80 percent, the proportion of cache misses rose from 10 percent to 20 percent, yet the cost increase typically far exceeded the magnitude of this numerical shift, since missed cache requests require complete prefill computation.[1] In pricing structures, cached input costs can be approximately one-tenth of the expense incurred by recalculation.[1] These dynamics underscore the criticality of infrastructure stability—assuming a workflow requires 20 consecutive model calls with a 99 percent success rate per call, the probability of completing the full chain successfully drops to approximately 82 percent.[1]
To address these measurement gaps, Qingcheng Extreme Intelligence developed the AI Ping evaluation platform, which covers more than 30 service providers and over 600 model services.[1] Recent market signals suggest gradual improvement: since 2026, certain domestic chip idle capacity has noticeably recovered, with some new chip models experiencing supply tightness within larger cluster deployments.[1]