谷歌推出了两款新型文本转语音模型——Gemini 3.8 Flash TTS和Gemini 3.8 Flash-Lite TTS1。其中Flash版本侧重创意应用,Flash-Lite版本则优化了效率以支持大规模部署1。这些模型能够根据自然语言提示生成自定义语音,支持超过100种语言和方言,涵盖墨西哥西班牙语、魁北克法语和苏格兰英语等区域变种1,并提供2000多个生产就绪的语音库1。
新模型具备多项先进功能。用户仅需提供30秒的音频样本即可实现语音复制,但该功能要求获得语音所有者的口头同意录音验证1。所有音频输出都配备SynthID水印标记,以识别AI生成的语音并防止虚假信息传播1。在Hume AI Voice Design Benchmark测试中,Flash版本得分71.4分排名第一,Flash和Flash-Lite在整体质量指数中分别位居第一和第二1。
该功能已在Gemini API、Google AI Studio和Gemini Notebook中开始推出1,旨在为创作者、开发者和企业提供更丰富、更具表现力的音频体验1。谷歌的合作伙伴包括Figma、HeyGen、Linguana、Wondercraft、99.co和Ollang1。
Google has introduced two new text-to-speech models designed to transform natural language prompts into custom voices 1. Gemini 3.8 Flash TTS targets creative applications, while Gemini 3.8 Flash-Lite TTS focuses on efficient, scalable deployment 1. Both models support over 100 languages and dialects, including regional variants such as Mexican Spanish, Quebec French, and Scottish English 1.
The models feature an extensive library of more than 2,000 production-ready voices 1. A voice cloning capability allows users to recreate vocal profiles using just 30 seconds of audio sample, though the feature requires verbal consent from the voice owner for verification purposes 1. In the Hume AI Voice Design Benchmark, Flash ranked first and Flash-Lite ranked second in overall quality, with Flash scoring 71.4 points 1. All audio output is embedded with SynthID watermarks to detect AI-generated speech and help prevent misinformation 1.
The new models are now rolling out across the Gemini API, Google AI Studio, and Gemini Notebook 1. Early partners utilizing the technology include Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang 1.
评论
还没有评论,欢迎留下第一条。