740M EmbeddingGemma 2 Maps Text, Image, Video and Audio
In this article
Google DeepMind has released EmbeddingGemma 2, an open-weights embedding model built on the Gemma 4 architecture that maps text, code, images, video, and audio into a shared 768-dimensional vector space. Released under the Apache 2.0 license, the 740-million-parameter model gives engineering teams one multimodal retrieval asset instead of separate text, vision, and audio encoders.
The model collapses several specialized encoders into a modular framework aimed at RAG and semantic search. It handles interleaved modalities inside an 8,192-token context window, so a text query can retrieve a photo or an audio trigger can surface a video clip without maintaining separate embedding clusters or tagging pipelines.
Modular architecture and footprint
The core is a 270-million-parameter text backbone: a 130-million-parameter transformer plus a 140-million-parameter embedder, covering more than 100 languages and code. Developers can attach a 170-million-parameter vision encoder, a 300-million-parameter audio encoder for raw 16 kHz mono audio, or both. Loading only the needed encoders sets the footprint at 270M parameters for text, 440M for text plus vision, 570M for text plus audio, and the full 740M for the complete multimodal pipeline. MarkTechPost reports that the full model is a 1.3GB Ollama download, while the 270M text-only variant takes 378MB. Google states that on a Pixel 11 Pro, quantized text-only weights occupy about 191MB of active RAM and the full multimodal configuration uses about 567MB; quantization-aware training targets INT4 and INT8.
Context and Matryoshka truncation
The model runs 24 layers with a 512-dimensional model space, 2,048-dimensional hidden layers, and a 262,144 vocabulary. It uses GQA/MQA attention with 4 heads, a 2:1 local-to-global KV-head ratio, and a 1,024-token sliding window. Activations are Gated FFN with GELU, with mean pooling and a 512-to-768 projection layer.
The 8,192-token budget holds roughly 5.5 minutes of audio, 58 video frames at 1 fps, or 29 images at standard token rates. A tighter 70-token vision budget fits roughly 114 images or frames.
Matryoshka Representation Learning lets developers truncate the native 768-dimensional outputs to 512, 256, or 128 dimensions. Moving to 128 dimensions cuts vector storage by up to 6x, but Google recommends reserving 128d mainly for text-only workloads because the MMEB score drops to 45.65 at that compression. At 256 dimensions, Google DeepMind reports the MTEB multilingual score slips only from 61.36 to 60.41.
Vendor-reported benchmarks
Google DeepMind reports leading scores among sub-1B multimodal embedders on MTEB Code and MAEB. At 768 full-precision dimensions, EmbeddingGemma 2 scores 61.36 on MTEB multilingual v2 and 78.68 on MTEB Code v1. The code result is a 9.92-point gain over EmbeddingGemma 1, roughly 14%. The maker also reports 64.64 on MIEB lite for images, 69.54 on MSEB retrieval for sound, 49.39 on MAEB audio, and 59.01 on MMEB v2. MarkTechPost notes that Alibaba's Qwen3-VL-Embedding-2B reports a 73.2 on its own MMEB-v2 run with about 2.7 times the parameters and no audio support.
| Model | Total Parameters | Supported Modalities | Context Window | Output Dimensions (MRL) | License |
|---|---|---|---|---|---|
| EmbeddingGemma 2 | 740M | Text, Code, Images, Video, Audio | 8,192 tokens | 768 (128/256/512) | Apache 2.0 |
| EmbeddingGemma 1 | 308M | Text, Code | 2K tokens | 768 (down to 128) | Gemma terms |
| Qwen3-VL-Embedding-2B | 2B | Text, Code, Images, Video | 32K tokens | Up to 2048 | Apache 2.0 |
AI Mastery analysis
Collapsing four modalities into one 768-dimensional space removes the orchestrator layer that previously aligned separate text, vision, and audio indexers. Because infrastructure rewrites not model weights drive 2026 ai gain, the larger operational win here is the smaller edge-deployable pipeline rather than the weights alone. Selective encoder loading reinforces that: inactive parameters stay out of RAM, which matters most on phones and laptops. The Google AI Edge team's 37.3 ms per image on a MacBook M5 Pro GPU, measured with a 70-token vision budget, is fast enough for interactive retrieval rather than batch indexing. Apache 2.0 licensing, ML Kit support for Android expected within weeks, and sentence-transformers v6.1.0+ compatibility make the model immediately usable on local hardware without cloud embedding APIs.
Sources
Frequently asked questions
How much memory does EmbeddingGemma 2 need on a phone?
Google reports about 191MB of active RAM for quantized text-only weights and about 567MB for the full multimodal model on a Pixel 11 Pro. Quantization-aware training targets INT4 and INT8 precision.
What changed between EmbeddingGemma 1 and EmbeddingGemma 2?
EmbeddingGemma 2 adds vision and audio encoders, moves the context window from 2K to 8,192 tokens, and raises MTEB Code v1 from 68.76 to 78.68. It also switches to an Apache 2.0 license.
Does EmbeddingGemma 2 support audio and video?
Yes. It maps text, images, video, and raw 16 kHz mono audio into one 768-dimensional space. Video inputs are sampled at 1 fps, with an 8,192-token window holding roughly 58 frames.
How does Matryoshka truncation reduce EmbeddingGemma 2 storage costs?
Truncating from 768 to 128 dimensions cuts vector storage by up to 6x. Google recommends 128d mainly for text workloads because the MMEB score drops to 45.65, while 256d keeps MTEB multilingual at 60.41 versus 61.36.
Can EmbeddingGemma 2 be used commercially?
Yes. The model is released under the Apache 2.0 license, and weights are available on Hugging Face and Kaggle with Ollama, GGUF, and LiteRT builds.
Related Reading
Perplexity's 9B Contextual Embedder Beats voyage-context-4 by 14.4 Points
Perplexity releases pplx-embed-v2-context-9b-preview, an open-weights 9B RAG embedder that scores 45.5% Answer Recall@10 and cuts storage 8× vs float32.
Qwen-Image-2.1: 7B Model Beats 32B FLUX 2 Max on Qwen's Benchmark
Alibaba's Qwen-Image-2.1 unifies image generation and editing in a 7B diffusion transformer, scoring 60.28 vs FLUX 2 Max's 55.33 at a quarter the parameters.
MiniCPM5-2B: 2.5B Model Beats Qwen3.5-4B With 53.9 Avg Across 34 Benchmarks
OpenBMB's MiniCPM5-2B averages 53.9 across 34 benchmarks, outscoring Qwen3.5-4B at 51.1 with roughly half the parameters and Apache 2.0 weights.