EmbeddingGemma 2 is Google DeepMind's new open embedding model, released on October 6, 2026. It turns text, code, images, video and audio into vectors in one shared 768-dimensional space, so a text query can find a matching photo, video clip or audio recording. It has 740 million parameters in total, but a text-only setup loads just 270 million of them, which makes it small enough for a laptop or phone. The weights are on the Hugging Face Hub under the Apache 2.0 licence, and the quickest way to use them is the Sentence Transformers library.

This guide covers what the model is, how to run it on your own machine, the settings that quietly decide whether your results are good or bad, and which local builds (GGUF, MLX, LiteRT) exist so far.

What EmbeddingGemma 2 is

An embedding model doesn't write text. It reads an input and returns a list of numbers (a vector) that captures its meaning, so similar inputs end up close together. You use those vectors for semantic search, retrieval-augmented generation (RAG), classification and clustering. EmbeddingGemma 2 is the successor to last year's text-only EmbeddingGemma, and according to the official model card it is built on the Gemma 4 architecture.

The model is modular. A 270M text part (a 130M transformer backbone plus a 140M embedder) is always loaded. A 170M vision encoder and a 300M audio encoder are optional, and you only load the ones you need.

FactEmbeddingGemma 2
Total parameters740M (text 270M, vision 170M, audio 300M)
InputsText (including code), images, video, audio, and mixes of them
Output768 dimensions, truncatable to 512, 256 or 128
Context window8,192 tokens shared across all inputs
Languages100+
LicenceApache 2.0
ReleasedOctober 6, 2026

Google's launch post says the first EmbeddingGemma passed 20 million downloads, and that the new context window is four times larger than its predecessor's. Since the weights are downloadable and the licence is permissive, it fits the usual meaning of an open-weight model; see open weight vs open source for what that does and doesn't cover.

How good is it?

The model card reports scores for the full-precision checkpoint. On multilingual text (MTEB multilingual v2) it scores 61.36, roughly level with the original EmbeddingGemma's 61.15. The big jump is code: MTEB Code rises from 68.76 to 78.68, which the launch post calls a 9.92-point gain. That makes the new model a natural pick for searching a local codebase or feeding a coding agent. It also posts new scores on image, video, visual-document and audio benchmarks, where version 1 had no support at all.

Treat these as the vendor's own numbers. Embedding quality depends heavily on your data, so test it on a sample of your own queries before you switch a production index over.

Run it locally with Sentence Transformers

The model card's quick start uses Sentence Transformers on top of Transformers. Install both:

pip install -U sentence-transformers transformers

Then embed a query and a document and compare them:

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("google/embeddinggemma-2")

query_emb = model.encode("What causes the northern lights?", prompt_name="SearchQuery")
doc_emb = model.encode("The northern lights are caused by charged particles from the sun.",
                       prompt_name="Document")
print(model.similarity(query_emb, doc_emb))

The first run downloads the weights from Hugging Face. After that everything runs on your machine, so the text, images and audio you embed never leave it. The repository isn't gated, so you don't need to accept a licence form or log in to download it.

Load only the encoders you need

If you're only embedding text, skip the vision and audio encoders to save memory. The model card shows how to do this by passing config_kwargs to Sentence Transformers:

What you embedconfig_kwargsEffective size
Text only{"vision_config": None, "audio_config": None}270M
Text and images{"audio_config": None}440M
Text and audio{"vision_config": None}570M
Everything{}740M

For example: SentenceTransformer("google/embeddinggemma-2", config_kwargs={"vision_config": None, "audio_config": None}). The card notes that other libraries handle this differently, so check their docs if you aren't using Sentence Transformers.

Use bfloat16 or float32, never float16

This is the setting that's easiest to get wrong. The model card warns that the model's activations go beyond what float16 can represent. In float16 it returns NaN or silently degraded embeddings instead of raising an error. Use bfloat16 on GPUs that support it (it halves memory compared with float32) and float32 everywhere else, including most CPUs:

import torch
dtype = torch.bfloat16 if torch.cuda.is_bf16_supported() else torch.float32
model = SentenceTransformer("google/embeddinggemma-2", model_kwargs={"torch_dtype": dtype})

Task prefixes: the quiet quality lever

EmbeddingGemma 2 was trained with short instructions in front of the text, and using the right one improves results. The model works without them, but less precisely. Prefixes apply to text only; images, video and audio go in without one.

  • Search and retrieval are asymmetric: queries get task: search result | query: ... and documents get title: ... | text: .... Use title: none when a document has no title.
  • Question answering, fact checking and code search follow the same pattern with their own query prefix, such as task: code retrieval | query: ....
  • Classification, clustering and sentence similarity are symmetric: every input gets the same prefix, such as task: clustering | query: ....

In Sentence Transformers, the prompt_name argument applies these for you (SearchQuery, QuestionAnswering, FactChecking, CodeRetrieval, Classification, Clustering, SentenceSimilarity, Document). One catch from the card: prompt_name="Document" always uses title: none, so if your documents have real titles, format them yourself, as in model.encode(f"title: {title} | text: {body}").

Shrink vectors with Matryoshka truncation

The model was trained with Matryoshka Representation Learning, so you can keep only the first 512, 256 or 128 numbers of each vector. That cuts vector storage by up to six times. The card's table shows multilingual MTEB going from 61.36 at 768 dimensions to 60.41 at 256 and 57.89 at 128. Its advice: quality is close to lossless down to 256, while 128 hurts multimodal quality a lot and suits text-only work best.

Two rules matter. Re-normalise the vector after truncating it, or cosine scores drift without any error. And queries and documents must use the same dimension. Sentence Transformers handles the first rule if you pass truncate_dim=256, normalize_embeddings=True to encode().

Images, video and audio

A single input can mix text with media. You mark where each item goes with placeholder tokens (<|image|>, <|video|>, <|audio|>) and pass the files under matching keys. The result is one vector you can compare against a plain text query. Everything shares the 8,192-token budget, and each kind of input uses it at a fixed rate:

  • An image costs 280 tokens by default, so about 29 fit.
  • Video is sampled at 1 frame per second, at 140 tokens per frame, so about 58 frames fit.
  • Audio costs 25 tokens per second, so roughly five and a half minutes fit. It should be mono at 16 kHz.

You can lower the vision token budget (anywhere from 70 to 1,120 tokens per image) to fit up to about 114 images or frames, trading detail for capacity.

GGUF, MLX and on-device builds

If you'd rather not run Python, several converted builds were already on Hugging Face within a day of release:

  • GGUF for llama.cpp: ggml-org/embeddinggemma-2-GGUF has the text model at BF16 (about 558 MB) and Q8_0 (about 310 MB), plus separate mmproj files for the vision and audio side (about 982 MB at BF16, 555 MB at Q8_0). Unsloth publishes a GGUF repo too. If the quant names are new to you, see GGUF quantization explained.
  • MLX for Apple Silicon: mlx-community has bf16, 8-bit, 4-bit, mxfp8, mxfp4 and nvfp4 conversions, for use with mlx-lm and the wider MLX ecosystem. Our MLX vs GGUF on Mac guide explains how to choose between them.
  • LiteRT for phones: litert-community publishes builds of the full 740M model, the 270M text-only model and the 440M text-and-vision model. The launch post says that, quantised on a Pixel 11 Pro, the model needs about 191 MB of active RAM for text alone and about 567 MB for everything.
  • ONNX: onnx-community has an ONNX conversion for ONNX-based runtimes.

These conversions are new. Check each repo's notes for which inputs it supports and which runtime version it needs, and test that its vectors match the reference model on a few samples. A different conversion can produce vectors that aren't compatible with the original, so embed your whole index with one build.

At the time of writing, Ollama's library lists the original embeddinggemma but not version 2.

Memory and hardware

This model is small next to the chat models covered in our VRAM guide. The full multimodal model at BF16 is about 1.5 GB of weights in the GGUF repo, and text-only Q8_0 is about 310 MB, so any recent laptop runs it on CPU. A GPU or Apple Silicon mostly helps when you're embedding large batches, such as a whole document archive or hours of audio.

Pairing it with a local chat model

Google pitches EmbeddingGemma 2 as the retrieval half of an on-device RAG pipeline, with a Gemma 4 model doing the answering. The two share a text tokenizer and audio encoder, which the launch post says lowers combined memory. It works with any local generator, though. You could retrieve with EmbeddingGemma 2 and answer with a model from our Qwen 3.8 local guide.

FAQ

Is EmbeddingGemma 2 free for commercial use?

Yes. It's released under Apache 2.0, and Google's model card adds that use must also follow the Gemma Prohibited Use Policy.

Can I use it for text only?

Yes. Disable the vision and audio encoders and only the 270M text part loads. The text model is always loaded, and the card's own example for text-only work drops the other two encoders.

Do I need to re-embed my EmbeddingGemma 1 index?

You should assume so. Vectors from two different models don't share a space, so mixing them in one index breaks similarity search. Re-embed your corpus with version 2 before switching queries over.

Why are my similarity scores NaN or strange?

The two usual causes are running in float16 (use bfloat16 or float32) and truncating vectors without re-normalising them. Using the wrong or missing task prefix also lowers quality, but less dramatically.

Does it run in Ollama?

Not yet as far as we can see: Ollama's library page lists only the original EmbeddingGemma. For a non-Python setup today, use the GGUF build with llama.cpp or the MLX builds on a Mac.