Cactus Compute团队发布了Whistle语音识别模型,文件体积仅为16.9 MB,可在CPU上运行且无任何依赖1。该模型支持转录、词级时间戳和语音嵌入三项核心功能,覆盖英语、德语、法语、西班牙语、意大利语、荷兰语和波兰语七种语言1。
Whistle在多项基准测试中的性能超越OpenAI的Whisper base模型,同时提供更快的处理速度和显著更小的体积——相比Whisper base的145.3 MB,体积缩小至16.9 MB1。在处理延迟方面,该模型处理5秒音频时首词延迟为5.9毫秒,处理30秒音频时为36.3毫秒1。
该模型已支持17个目标平台的部署,包括macOS、Linux、Android、iOS、watchOS、Windows on ARM、RISC-V、MIPS、浏览器以及WASI组件1。开发团队通过比对音频校验和与说话人ID确保评估所用的86,174个语句未出现在模型的训练或验证数据中1。
The Cactus Compute team has released Whistle, a speech-to-text model that weighs just 16.9 MB and runs on CPU without external dependencies.1 The model supports transcription, word-level timestamps, and speech embeddings across seven languages: English, German, French, Spanish, Italian, Dutch, and Polish.1
Whistle demonstrates superior performance compared to OpenAI's Whisper base model across multiple benchmarks while delivering faster processing speeds and a significantly smaller footprint—Whisper base requires 145.3 MB by comparison.1 The model achieves notably low latency, producing the first word in 5.9 milliseconds for five-second audio and 36.3 milliseconds for thirty-second audio.1 Performance was evaluated on a dataset of 86,174 utterances, with validation protocols ensuring test audio did not appear in training or validation data by cross-referencing audio checksums and speaker IDs.1
The model is optimized for deployment across 17 target platforms, including macOS, Linux, Android, iOS, watchOS, Windows on ARM, RISC-V, MIPS, web browsers, and WASI components.1
评论
还没有评论,欢迎留下第一条。