Nari Labs公开了其开源的Qwen3-TTS和Qwen3-ASR语音模型及推理引擎1。在Coval基准测试中,这两款模型的综合表现超越了11Labs、Cartesia等闭源方案1。其中Qwen3-TTS在准确度指标上排名第一,延迟排名第二,同时成本最低1;Qwen3-ASR则实现了最低延迟的同时保持第二优的准确度,仅与最优准确度相差0.1%,成本处于次低水平1。
该推理引擎在每秒10次请求的并发下可实现低于50毫秒的延迟1。Nari Labs已将开源推理引擎代码在GitHub上发布1。这是继去年发布首个开源自然对话TTS模型Dia之后,该团队在语音技术领域的又一次更新1。
Nari Labs has unveiled its open-source Qwen3-TTS and Qwen3-ASR speech models along with an inference engine, demonstrating superior performance against closed-source alternatives in the Coval benchmark tests1. The Qwen3-TTS model ranks first in accuracy with the lowest word error rate (WER) and second in latency when compared against eleven competing models including 11Labs and Cartesia, while maintaining the lowest operational cost1. The Qwen3-ASR system achieves the fastest latency performance with accuracy ranking second, trailing by only 0.1 percent, and offers the second-lowest inference cost1.
The inference engine delivers sub-50 millisecond latency at 10 requests per second, enabling efficient real-time speech processing1. Nari Labs has made the open-source inference engine publicly available on GitHub at https://github.com/nari-labs/nari-qwen3-tts[1](#source-1). The company previously introduced Dia, the first open-source natural conversation TTS model1.
评论
还没有评论,欢迎留下第一条。