SpaceX AI推出新一代语音模型Grok Voice Think Fast 2.0,在多项基准测试中取得领先成绩[1]。该模型在Tau Voice基准测试中以56.5%的得分排名第一,击败OpenAI、Google等竞争对手[1]。在Artificial Analysis的Speech to Speech Index测试中,高推理能力版本得分达82.9%,仅次于千问Qwen Audio 3.0 Realtime Plus的84.1%[1]。
Think Fast 2.0在响应速度和成本效率上均实现显著提升[1]。平均首音频响应时间仅为0.70秒,是榜单前五名中唯一低于1秒的模型,相比前代的1.25秒大幅改进[1]。定价方面,该模型每分钟收费0.08美元(约每小时4.80美元),仅为GPT-Realtime-2.1 High版本的45%[1]。在转写准确率上,Think Fast 2.0相比前代提升约1.4倍,相比Deepgram Nova 3和ElevenLabs Scribe v2的提升幅度为1.5至2倍[1]。此外,该模型的中位数推理Token消耗仅为前代的约40%,综合得分相比上一代的75.7%提升至82.9%,增幅达7.2个百分点[1]。
Elon Musk's SpaceX AI has unveiled Grok Voice Think Fast 2.0, a next-generation speech-to-speech model that has achieved top rankings in industry benchmarks [1]. The model scored 56.5% on the Tau Voice benchmark test, securing the first-place position ahead of competitors including OpenAI and Google [1]. Additionally, in the Artificial Analysis Speech to Speech Index, the high-reasoning version of the model achieved a score of 82.9%, placing it just behind Qwen Audio 3.0 Realtime Plus at 84.1% [1].
The model delivers significant performance improvements in both speed and accuracy [1]. With an average first audio response time of just 0.70 seconds, Grok Voice Think Fast 2.0 is the only model in the top five rankings to remain below the one-second threshold—a substantial improvement from the previous generation's 1.25 seconds [1]. The transcription accuracy has increased by approximately 1.4 times compared to Think Fast 1.0, and by 1.5 to 2 times compared to competitors Deepgram Nova 3 and ElevenLabs Scribe v2 [1]. The model also demonstrates more efficient token consumption, with median reasoning token usage dropping to roughly 40% of the previous generation [1].
Pricing represents a competitive advantage for developers [1]. At $0.08 per minute for audio input—equivalent to $4.80 per hour—Grok Voice Think Fast 2.0 costs approximately 45% of OpenAI's GPT-Realtime-2.1 High offering [1]. The comprehensive performance score of 82.9% marks a 7.2 percentage point improvement over the previous model's 75.7% [1].