谷歌推出了EmbeddingGemma 2,一款开源轻量级多模态嵌入模型,包含7.4亿参数,在Apache 2.0许可证下发布12。该模型支持将文本、代码、图像、视频和音频映射到统一的嵌入空间12,优化用于设备端推理,前代版本已获得超过2000万次下载1。
在性能方面,EmbeddingGemma 2的代码性能在MTEB Code基准上从68.76提升至78.68,增幅达9.92分1。模型的上下文窗口从4K扩展到8K token,较仅文本版本扩大4倍12,可处理5.5分钟音频、29张图像或58个视频帧1。文本仅模式需要约191MB主动RAM,完整多模态模型需要约567MB1。
EmbeddingGemma 2支持从768维动态截断至512、256或128维1,这项功能可将存储需求减少6倍1,通过Matryoshka表示学习实现向量存储压缩1。
Google has introduced EmbeddingGemma 2, an open-source lightweight multimodal embedding model designed to map text, code, images, video, and audio into a unified embedding space 12. The model contains 740 million parameters and is released under the Apache 2.0 license, optimized for on-device inference 12.
The new model represents a significant advancement over its predecessor, which has been downloaded more than 20 million times 1. EmbeddingGemma 2 expands its context window to 8,000 tokens—a fourfold increase from the text-only version—enabling it to process up to 5.5 minutes of audio, 29 images, or 58 video frames in a single pass 1. Performance improvements are substantial, with code benchmark scores on MTEB rising from 68.76 to 78.68, a gain of 9.92 points 1.
The model offers flexible deployment options tailored to different hardware constraints. The text-only variant requires approximately 191 megabytes of active RAM, while the full multimodal model requires around 567 megabytes 1. Vector dimensions can be dynamically truncated from 768 down to 512, 256, or 128 dimensions through Matryoshka representation learning, reducing storage requirements by up to six times 1. These capabilities make EmbeddingGemma 2 suitable for resource-constrained environments while maintaining strong performance across multiple benchmark tests 12.
评论
还没有评论,欢迎留下第一条。