Wispr Advanced Interfaces Lab推出了Canto语音模型,专门面向实际听写应用场景设计1。该模型在由2,300多名独特说话者组成、包含10小时英语听写数据的真实评估集上实现了最低的词错率(WER),超越Google、OpenAI、AssemblyAI和Deepgram等竞争对手的模型1。在3小时挑战评估集的实时转录模型中,Canto排名最低WER,仅次于Gemini 3.1 Pro(但后者不适合实时低延迟应用)1。同时,Canto在LibriSpeech公开基准上与竞争对手并列最低WER1。
Canto采用两阶段训练方式,结合监督微调和强化学习方法1。具体来说,该模型使用Group Relative Policy Optimization(GRPO)技术,通过比较候选转录本的相对奖励来指导训练过程1。这种训练方式使其能在背景噪音、低音量和短听写等具有挑战性的条件下保持良好性能1。
后续模型将在Canto规模的十倍以上进行训练,改进目标包括增强困难音频条件下的识别能力、提升多说话者识别效果以及拓展多语言支持1。
Wispr Advanced Interfaces Lab has introduced Canto, a speech recognition model designed specifically for real-world dictation applications.1 The model achieved the lowest word error rate (WER) when evaluated on authentic user dictation data, outperforming competing systems from Google, OpenAI, AssemblyAI, and Deepgram.1 Canto was evaluated on a real assessment set comprising over 10 hours of English dictation from more than 2,300 unique speakers, where it demonstrated superior performance.1
The model employs a two-stage training approach combining supervised fine-tuning with Group Relative Policy Optimization (GRPO), a reinforcement learning technique that guides training by comparing relative rewards of candidate transcriptions.1 This approach enables Canto to maintain robust performance under challenging audio conditions, including background noise, low volume, and brief dictation segments.1 On a three-hour challenge evaluation set, Canto ranked lowest in WER among real-time transcription models, second only to Gemini 3.1 Pro, which is unsuitable for real-time low-latency applications.1 The model also achieved competitive performance on the LibriSpeech public benchmark, matching the lowest WER of competing systems.1
Looking ahead, Wispr plans to scale subsequent versions of Canto to ten times or larger, with objectives including improved recognition of difficult audio conditions, multi-speaker identification, and multilingual support.1
评论
还没有评论,欢迎留下第一条。