Google DeepMind 发布 EmbeddingGemma 2 多模态嵌入模型
Introducing EmbeddingGemma 2: A best-in-class open model for natively multimodal embeddings | Google
Google DeepMind 发布开源多模态嵌入模型 EmbeddingGemma 2,可将文本(含代码)、图像、视频和音频及其组合映射到统一的 768 维向量空间。模型总参数 740M,由 270M 文本模型与 170M 视觉编码器、300M 音频编码器组成,面向手机和笔记本等消费级硬件,用于端侧搜索、RAG、分类和聚类等低延迟语义表示场景。模型已发布在 Hugging Face。
| https://huggingface.co/google/embeddinggemma-2 EmbeddingGemma 2 is an open multimodal embedding model built by Google DeepMind which maps text (incl. code), images, video, and audio inputs—and combinations thereof—into a single, unified 768-dimensional vector space. The model has 740M total parameters, combining a 270M parameter text model with modular vision (170M) and audio (300M) encoders. Designed to run on consumer hardware such as mobile devices and laptops, EmbeddingGemma 2 delivers low-latency semantic representations for on-device applications, like search, retrieval-augmented generation (RAG), classification, and clustering. submitted by /u/Recoil42[link] [留言] |
来源:r/LocalLLaMA · reddit.com