The Spectrum Dispatch News

technology

Google launches EmbeddingGemma 2, a lightweight multimodal embedding model

The new model unifies text, code, images, video, and audio in a shared embedding space, enabling on-device search and retrieval with minimal resource requirements.

Google launches EmbeddingGemma 2, a lightweight multimodal embedding model

Google has released EmbeddingGemma 2, an expanded version of its lightweight embedding model that now supports multiple data types beyond text. According to Google, the original EmbeddingGemma achieved more than 20 million downloads and was used by developers to build on-device search tools and privacy-first retrieval augmented generation (RAG) pipelines.

Google launches EmbeddingGemma 2, a lightweight multimodal embedding model

EmbeddingGemma 2 extends this capability to unify code, images, video, and audio in a shared embedding space. Built on the Gemma 4 architecture with 740 million parameters, the model is released under an Apache 2.0 license and is optimized for on-device inference.

The model offers several key capabilities. It achieves leading scores among sub-1 billion parameter multimodal embedders on benchmarks like MTEB (Massive Text Embedding Benchmark) Code and MAEB (Massive Audio Embedding Benchmark). According to Google, it showed a 9.92-point improvement on code performance compared to the original EmbeddingGemma, scoring 78.68 on MTEB Code versus 68.76 previously.

EmbeddingGemma 2 is designed with modularity in mind. For text-only workloads, it requires as little as 270 million parameters, with optional vision (170M) and audio (300M) encoders available for full multimodal support. Using Matryoshka Representation Learning, developers can dynamically truncate output vectors from 768 dimensions down to 128, achieving up to 6x storage reduction for local vector databases.

The model supports an 8K token context window—four times larger than the original EmbeddingGemma—allowing it to process up to 5.5 minutes of audio, 29 images, or 58 video frames directly on local hardware. According to Google, on a Google Pixel 11 Pro with quantization, EmbeddingGemma 2 requires approximately 191MB active RAM for text-only weights and 567MB for the full multimodal model.

When paired with generative models like Gemma 4, EmbeddingGemma 2 enables on-device RAG pipelines that process multimodal data while maintaining data privacy and reducing latency. Since both models share a text tokenizer and audio encoder, they can run together with a lower combined memory footprint.

Google has made EmbeddingGemma 2 available for download on Hugging Face and Kaggle, with Gemini Enterprise Agent Platform Model Garden availability coming soon. The model can be deployed using tools like Google AI Edge MediaPipe, LiteRT, transformers.js, and other frameworks including vLLM, llama.cpp, and Ollama.

Key facts

  • EmbeddingGemma 2 has 740 million parameters and supports text, code, images, video, and audio in a single model
  • The original EmbeddingGemma achieved more than 20 million downloads since its release last year
  • Code performance improved 9.92 points on MTEB Code benchmark (78.68 vs 68.76)
  • The model features an 8K token context window, 4x larger than the original
  • On-device deployment requires ~191MB RAM for text-only and ~567MB for full multimodal on Google Pixel 11 Pro
  • Storage can be reduced up to 6x through dynamic vector truncation

Sources

← All posts