Changes confirmed medium confidence

Google Releases EmbeddingGemma 2 for On-Device Multimodal Search

The open 740-million-parameter model puts text, code, images, video and audio into one retrieval space, while its efficiency and benchmark claims still come from Google's own launch testing.

Edited by Tyronne Panaino

Google launched EmbeddingGemma 2 on October 6, 2026 as an open, on-device embedding model that maps text, code, images, video and audio into a shared vector space. The release matters to developers building local search, retrieval-augmented generation and routing because one model can compare meaning across several media types without sending every item to a hosted embedding service.

The model is built on the Gemma 4 architecture, carries an Apache 2.0 licence and has 740 million parameters in its full configuration. Google presents the release as a step beyond the first EmbeddingGemma, which focused on text, but the important practical delta is broader than adding another input type: the same embedding space can connect a text query to a video moment, an audio query to stored media or code to a semantically related passage.

A modular model for different device budgets

EmbeddingGemma 2 is modular rather than all-or-nothing. Google says a text-only workload can use 270 million parameters, while optional vision and audio encoders add 170 million and 300 million parameters respectively. That lets an application omit media components it does not need instead of carrying the full multimodal configuration.

The launch article reports an 8K-token context window. Google translates that capacity into several modality-specific examples: as much as 5.5 minutes of audio, 29 images or 58 video frames, including interleaved combinations. Those are input ceilings described by the vendor, not evidence that every device will process the largest supported input with acceptable latency or energy use.

Google also reports memory figures from one reference device. With quantization on a Pixel 11 Pro, it says text-only weights can require about 191 MB of active RAM and the full multimodal model about 567 MB. These measurements make the on-device goal concrete, but they do not establish equivalent performance across older phones, browsers, laptops or embedded hardware.

Adjustable vectors change the storage calculation

The model uses Matryoshka Representation Learning so developers can shorten its 768-dimensional output vectors to 512, 256 or 128 dimensions. Google says this can reduce local vector-database storage and memory use by as much as six times.

That flexibility creates an application-level trade-off rather than a universal saving. Shorter vectors use less space and can lower retrieval costs, while each team still needs to measure whether reduced dimensions preserve enough ranking quality for its own corpus. The fetched launch article does not provide a workload-independent threshold at which a smaller vector stops being accurate enough.

The same caveat applies to Google's benchmark claims. The company reports that MTEB Code rose from 68.76 for the earlier model to 78.68 for EmbeddingGemma 2, a 9.92-point increase, and says the model leads sub-one-billion-parameter multimodal embedders across named evaluations. Those results are useful first-party evidence about the release, but the fetched material does not include an independent reproduction or show how the scores transfer to a specific private dataset.

What developers can use now

Google says the model weights are available through Hugging Face and Kaggle. The launch also lists deployment paths through MediaPipe, LiteRT, browser tooling and several common inference frameworks. Availability in the Gemini Enterprise Agent Platform Model Garden is described as coming later, so it should not be treated as part of the current release.

The most immediate use cases are local semantic search, media retrieval, classification and routing. Google also describes pairing EmbeddingGemma 2 with Gemma 4 for on-device retrieval-augmented generation. Keeping embeddings local can reduce how much raw media leaves a device, but privacy still depends on the rest of the application: a later generative step, analytics layer or synchronization service can reintroduce data transfer.

Evidence quality and the next checkpoints

The release is confirmed by one detailed first-party Google announcement. Internal confidence is medium because the architecture, licence, availability and reported measurements are official, while the quality, memory and storage comparisons remain vendor-produced. No independently fetched test establishes latency, battery impact, retrieval quality across languages and modalities, or reliability on a broad range of devices.

The next useful checkpoints are independent retrieval evaluations, device-level latency and energy measurements, and reports from teams deploying the modular encoders with different vector sizes. Those results will show whether the model's cross-modal convenience outweighs any accuracy or operating-cost trade-offs in production.

Status

Confirmed. Google has released the model weights and documented the architecture, supported modalities, modular components and reported measurements. The comparative performance and efficiency claims remain first-party evidence.

Sources

Update note: Last reviewed 2026-10-07. We will revise this post if Google changes availability or independent device and retrieval evaluations materially alter the evidence.

Sources

Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.

More Changes coverage