Google 发布了 Gemini 3.5 Transcribe 语音转文字模型,官方称其为最精确的语音识别模型,可将原始音频直接转换为准确且格式化的文本 1。开发者可通过 Google AI Studio 的 Gemini API 和 Gemini Enterprise Agent Platform 使用该模型 1。该模型提供两种接口:用于实时流式处理的 gemini-3.5-transcribe-live,以及用于预录音频处理的 gemini-3.5-transcribe 1。
在性能方面,据 Artificial Analysis 测量,该模型在流式场景的平均词错误率(WER)为 4.0%,非流式场景为 2.6% 1。在 FLEURS 基准测试中,其流式与非流式 WER 分别为 5.50% 和 5.04% 1。相比前代 Chirp 3,其最终转录时间缩短了 70% 1。该模型支持超过 85 种语言的自动检测与转录,且预录音频支持最多 3 位说话人的多说话人识别与时间戳 1。目前,该模型已集成至 Gboard(Android Rambler 功能)、Gemini macOS 应用和 Google Antigravity,并即将登陆 Chrome 1。
Google has introduced Gemini 3.5 Transcribe, a new speech-to-text model that the company claims is its most accurate speech recognition tool to date 1. The system is designed to convert raw audio directly into accurate and formatted text 1. It is currently available to developers through the Gemini API, providing two distinct interfaces: a real-time streaming endpoint named gemini-3.5-transcribe-live and a pre-recorded audio endpoint called gemini-3.5-transcribe 1.
Regarding performance, measurements conducted by Artificial Analysis show that the model achieves an average Word Error Rate (WER) of 4.0% in streaming scenarios and 2.6% in non-streaming scenarios 1. On the FLEURS benchmark, it recorded a streaming WER of 5.50% and a non-streaming WER of 5.04% 1. Compared to the previous generation, Chirp 3, Gemini 3.5 Transcribe reduces final transcription time by 70% 1. The model also supports automatic detection and transcription for more than 85 languages 1. Furthermore, the pre-recorded audio interface includes multi-speaker recognition and timestamps for up to three speakers 1.
Developers can access the new model via the Gemini API in Google AI Studio as well as the Gemini Enterprise Agent Platform 1. Beyond developer tools, Gemini 3.5 Transcribe has already been integrated into several Google products, including Gboard with the Android Rambler feature, the Gemini macOS application, and Google Antigravity 1. The technology is also scheduled to arrive in the Chrome browser soon 1.
评论
还没有评论,欢迎留下第一条。