740M EmbeddingGemma 2 Maps Text, Image, Video and Audio

October 6, 2026 • news
Open WeightsMultimodalEmbeddingsGoogle DeepMindOn-Device AI

Google DeepMind has released EmbeddingGemma 2, an open-weights embedding model built on the Gemma 4 architecture that maps text, code, images, video, and audio into a shared 768-dimensional vector space. Released under the Apache 2.0 license, the 740-million-parameter model gives engineering teams one multimodal retrieval asset instead of separate text, vision, and audio encoders.

The model collapses several specialized encoders into a modular framework aimed at RAG and semantic search. It handles interleaved modalities inside an 8,192-token context window, so a text query can retrieve a photo or an audio trigger can surface a video clip without maintaining separate embedding clusters or tagging pipelines.

Modular architecture and footprint

The core is a 270-million-parameter text backbone: a 130-million-parameter transformer plus a 140-million-parameter embedder, covering more than 100 languages and code. Developers can attach a 170-million-parameter vision encoder, a 300-million-parameter audio encoder for raw 16 kHz mono audio, or both. Loading only the needed encoders sets the footprint at 270M parameters for text, 440M for text plus vision, 570M for text plus audio, and the full 740M for the complete multimodal pipeline. MarkTechPost reports that the full model is a 1.3GB Ollama download, while the 270M text-only variant takes 378MB. Google states that on a Pixel 11 Pro, quantized text-only weights occupy about 191MB of active RAM and the full multimodal configuration uses about 567MB; quantization-aware training targets INT4 and INT8.

Context and Matryoshka truncation

The model runs 24 layers with a 512-dimensional model space, 2,048-dimensional hidden layers, and a 262,144 vocabulary. It uses GQA/MQA attention with 4 heads, a 2:1 local-to-global KV-head ratio, and a 1,024-token sliding window. Activations are Gated FFN with GELU, with mean pooling and a 512-to-768 projection layer.

The 8,192-token budget holds roughly 5.5 minutes of audio, 58 video frames at 1 fps, or 29 images at standard token rates. A tighter 70-token vision budget fits roughly 114 images or frames.

Matryoshka Representation Learning lets developers truncate the native 768-dimensional outputs to 512, 256, or 128 dimensions. Moving to 128 dimensions cuts vector storage by up to 6x, but Google recommends reserving 128d mainly for text-only workloads because the MMEB score drops to 45.65 at that compression. At 256 dimensions, Google DeepMind reports the MTEB multilingual score slips only from 61.36 to 60.41.

Vendor-reported benchmarks

Google DeepMind reports leading scores among sub-1B multimodal embedders on MTEB Code and MAEB. At 768 full-precision dimensions, EmbeddingGemma 2 scores 61.36 on MTEB multilingual v2 and 78.68 on MTEB Code v1. The code result is a 9.92-point gain over EmbeddingGemma 1, roughly 14%. The maker also reports 64.64 on MIEB lite for images, 69.54 on MSEB retrieval for sound, 49.39 on MAEB audio, and 59.01 on MMEB v2. MarkTechPost notes that Alibaba's Qwen3-VL-Embedding-2B reports a 73.2 on its own MMEB-v2 run with about 2.7 times the parameters and no audio support.

Model Total Parameters Supported Modalities Context Window Output Dimensions (MRL) License
EmbeddingGemma 2 740M Text, Code, Images, Video, Audio 8,192 tokens 768 (128/256/512) Apache 2.0
EmbeddingGemma 1 308M Text, Code 2K tokens 768 (down to 128) Gemma terms
Qwen3-VL-Embedding-2B 2B Text, Code, Images, Video 32K tokens Up to 2048 Apache 2.0

AI Mastery analysis

Collapsing four modalities into one 768-dimensional space removes the orchestrator layer that previously aligned separate text, vision, and audio indexers. Because infrastructure rewrites not model weights drive 2026 ai gain, the larger operational win here is the smaller edge-deployable pipeline rather than the weights alone. Selective encoder loading reinforces that: inactive parameters stay out of RAM, which matters most on phones and laptops. The Google AI Edge team's 37.3 ms per image on a MacBook M5 Pro GPU, measured with a 70-token vision budget, is fast enough for interactive retrieval rather than batch indexing. Apache 2.0 licensing, ML Kit support for Android expected within weeks, and sentence-transformers v6.1.0+ compatibility make the model immediately usable on local hardware without cloud embedding APIs.

Sources

Frequently asked questions

How much memory does EmbeddingGemma 2 need on a phone?

Google reports about 191MB of active RAM for quantized text-only weights and about 567MB for the full multimodal model on a Pixel 11 Pro. Quantization-aware training targets INT4 and INT8 precision.

What changed between EmbeddingGemma 1 and EmbeddingGemma 2?

EmbeddingGemma 2 adds vision and audio encoders, moves the context window from 2K to 8,192 tokens, and raises MTEB Code v1 from 68.76 to 78.68. It also switches to an Apache 2.0 license.

Does EmbeddingGemma 2 support audio and video?

Yes. It maps text, images, video, and raw 16 kHz mono audio into one 768-dimensional space. Video inputs are sampled at 1 fps, with an 8,192-token window holding roughly 58 frames.

How does Matryoshka truncation reduce EmbeddingGemma 2 storage costs?

Truncating from 768 to 128 dimensions cuts vector storage by up to 6x. Google recommends 128d mainly for text workloads because the MMEB score drops to 45.65, while 256d keeps MTEB multilingual at 60.41 versus 61.36.

Can EmbeddingGemma 2 be used commercially?

Yes. The model is released under the Apache 2.0 license, and weights are available on Hugging Face and Kaggle with Ollama, GGUF, and LiteRT builds.

Free interactive tools for the decisions this piece raises.

Related Reading