OpenAI于7月29日宣布,GPT-5.6 Sol模型已被部署到生产环境,可自主优化运行系统[1]。该模型通过独立重写用Triton和Gluon编程语言编写的生产内核,并在GPU编程、负载均衡、推测解码等多个环节进行优化,使OpenAI的端到端服务成本降低了20%,token生成效率提升超过15%[1][2]。
GPT-5.6 Sol执行的优化任务包括分析生产流量、重写生产Kernel、优化推测解码系统以及搜索最优部署参数[2]。该模型还针对draft模型设计并跑了数百次架构实验[1]。值得注意的是,人类仍掌控优化目标、工具权限、评测指标和代码发布权[2]。
内部数据反映了GPT-5.6应用规模的快速增长。在内部测试期间,每名活跃研究人员的日均token输出超过GPT-5.5峰值的两倍[1],过去半年内部编程推理的研究算力占比增长了100倍,智能体token用量增长了约22倍[1]。此外,GPT-5.6在RSI Index上比GPT-5.5高出16.2分[1]。
OpenAI has announced that GPT-5.6 Sol, its latest AI model, has successfully optimized the company's production systems by autonomously rewriting and refining core infrastructure components. [1][2] On July 29, the company disclosed that the model independently redesigned GPU kernels and the inference stack using Triton and Gluon programming languages, focusing on GPU programming, load balancing, and speculative decoding across multiple operational layers. [1]
The results of this self-optimization effort are substantial. [1][2] End-to-end service costs have decreased by 20%, while token generation efficiency has improved by over 15%. [1][2] During the internal testing phase, each active researcher produced daily token output more than double the peak values achieved with GPT-5.5. [1][2] The optimization effort encompassed four core engineering tasks: analyzing production traffic, rewriting production kernels, enhancing speculative decoding systems, and searching for optimal deployment parameters. [2]
Despite the model's expanded autonomy, humans retain critical oversight mechanisms. [2] OpenAI personnel maintain control over optimization objectives, tool access, evaluation metrics, and code deployment authority. [2] Over the past six months, the company's internal programming reasoning research computing allocation has surged 100-fold, while internal agent token consumption has grown approximately 22-fold. [1][2]