AI summarized from verified sources
Run image, audio, and video search on-device with one model
One lightweight model enables offline multimodal search and RAG across media types, improving privacy and speed.
SOURCE CHECK
2 sources
Sources
Key Points
- 1Unified embeddings for text, images, audio, video
- 2740M parameters optimized for on-device use
- 3Apache 2.0 license for commercial use and fine-tuning
Google DeepMind open-released EmbeddingGemma 2 on October 6, 2026. This 740M-parameter lightweight multimodal embedding model maps text, images, audio, and video into a unified 768-dimensional space. Licensed under Apache 2.0 and available on Hugging Face, it excels at on-device RAG and cross-modal search.
Key points
EmbeddingGemma 2 is a 740M-parameter model based on Gemma 4. It uses a modular 270M text/code core plus optional 170M vision and 300M audio encoders, supporting 8K context for long documents, video, and audio.
Impact
Developers can easily build privacy-first on-device RAG and cross-modal search. Weights are already available on Hugging Face and Kaggle.
What changed
Google DeepMind open-released EmbeddingGemma 2 on October 6, 2026. This 740M-parameter lightweight multimodal embedding model maps text, images, audio, and video into a unified 768-dimensional space. Licensed under Apache 2.0 and available on Hugging Face, it excels at on-device RAG and cross-modal search.
Briefs that include this news
Use daily, weekly, and monthly briefs to understand the surrounding context.