OpenAI与芯片制造商Cerebras合作推出Ultrafast模式,使GPT-5.6 Sol模型的处理速度实现显著提升[1][2]。该模式可达每秒750个输出令牌的速度[1][2],相较于此前的性能水平提升14倍[2]。在Humanity's Last Exam基准测试中,GPT-5.6 Sol Ultrafast完成2500道题仅需11小时11分钟,较Claude Fable 5的78小时27分钟快7倍[1]。同时在GDP-Val基准测试中实现了5.6倍的端到端加速,未产生质量损失[1]。
这一突破由Cerebras的Wafer-Scale Engine架构驱动,该芯片集成44 GB SRAM[1]。OpenAI表示,"直到现在,获得实时速度通常意味着选择更小或更专业化的模型。Ultrafast指向了一个新方向的进展:每秒完成更多有用的工作"[2]。该服务目前以预览版形式向部分客户限制开放,将逐步扩大访问权限[1]。该模式适用于事件响应、客户服务和支持、金融市场分析以及电子商务等企业工作流[2]。
OpenAI has introduced Ultrafast, a new inference mode that accelerates its flagship GPT-5.6 Sol model to unprecedented speeds.[2] The service, developed in collaboration with chip manufacturer Cerebras, achieves output speeds of 750 tokens per second.[1][2] Currently available as a limited preview to select customers, Ultrafast represents a significant leap in real-time AI processing capability.
The performance gains are substantial across multiple benchmarks. On the Humanity's Last Exam evaluation, GPT-5.6 Sol Ultrafast completed 2,500 questions in 11 hours and 11 minutes—seven times faster than Claude Fable 5, which required 78 hours and 27 minutes.[1] Additionally, the model demonstrated 5.6x end-to-end acceleration on the GDP-Val benchmark without compromising output quality.[1] The breakthrough is powered by Cerebras' Wafer-Scale Engine architecture, which integrates 44 GB of SRAM on a single chip.[1]
OpenAI framed the capability as opening a new direction for enterprise applications. "Until now, getting real-time speed typically meant choosing a smaller or more specialized model. Ultrafast points to progress in a new direction: more useful work per second," the company stated.[2] The service is designed for use cases including event response, customer service, financial market analysis, and e-commerce workflows.[2] Access is expected to expand gradually beyond the current preview phase.