Google DeepMind launches embeddinggemma 2, a multimodal embedding model for on-device applications
Google DeepMind has launched EmbeddingGemma 2, a 740M parameter open-source multimodal embedding model designed for on-device search, classification, and privacy-first RAG applications. The model, built on the Gemma 4 architecture, supports text, code, images, video, and audio, embedding them into a 768-dimensional space. It features an 8K token context window and is available under an Apache 2.0 license. Weights are now live on Hugging Face and Kaggle, with Oll, llama.cpp, and LiteRT builds available.
EmbeddingGemma 2 offers modular configurations, allowing developers to choose between 270M, 440M, 570M, or the full 740M parameters. The model achieves leading scores on MTEB Code and MAEB benchmarks, with full-precision results at 768 dimensions. Quantization reduces memory usage, with text-only weights requiring about 191MB on a Pixel 11 Pro. The model also supports Matryoshka Representation Learning, enabling dimension reduction for storage efficiency.
The model runs on multiple frameworks, including sentence-transformers, Transformers, vLLM, and MediaPipe. It is compatible with Qdrant for vector storage and Unsloth for fine-tuning. Google DeepMind also announced upcoming ML Kit support for Android with NPU acceleration.