Google Open-Sources 740M-Parameter Multimodal Embedding Model

On October 6, Google released EmbeddingGemma 2, an open-weight multimodal embedding model. Rather than handling only one type of input, it maps text, code, images, video frames and audio into the same shared vector space, which means a search across very different kinds of content can be done with one model. The explicit goal is on-device search on phones and laptops, where the data being indexed never has to leave the device.

The model is built on Gemma 4. The full version has about 740 million parameters in total, made up of three separate pieces that can be mixed and matched depending on the use case: a 270-million-parameter text backbone that handles pure text and code, plus an optional 170-million-parameter vision encoder for images and video frames and a separate 300-million-parameter audio encoder. If an application only needs text retrieval, installing the backbone alone is enough and keeps the footprint small. The whole thing is licensed under Apache 2.0, so it's free to use commercially as well as to modify.

Specs

The context length is 8K tokens. By Google's own conversion, a single input can fit roughly 5.5 minutes of audio, 29 images or 58 video frames.

The output is a 768-dimensional vector, which developers can truncate down to 512, 256 or 128 dimensions depending on how much precision they're willing to trade away. This relies on a technique called Matryoshka representation learning, where cutting the vector shorter causes a relatively small loss in accuracy. Google says that, at most, local storage and indexing can shrink to roughly one-sixth of the original footprint this way — though a separate part of the developer blog puts the maximum savings at up to 8x, so the two figures don't entirely line up.

On-device memory footprint is the headline number in this release. Quantized and running on a Pixel 11 Pro, the text-only version uses about 191MB of memory at runtime, which rises to about 567MB once every modality — text, image, audio and video — is switched on. On the M5 Pro GPU inside a MacBook, processing a single image takes about 37.3 milliseconds, or roughly 27 images every second. The model is distributed in both INT4 and INT8 quantized versions, giving developers a choice between size and precision.

On benchmarks, the code retrieval benchmark MTEB Code scored 78.68, up 9.92 points from the first generation's 68.76. On the MIEB Lite image embedding benchmark and the MAEB audio embedding benchmark, Google says it leads models of similar size, though the announcement doesn't list specific competitors or scores. Multilingual text retrieval is on par with the first generation.

A few companion demos

The Google AI Edge Gallery app, which is available on both Android and iOS, has added two new demos built on top of the model: instant media search, which lets someone find a photo in their gallery just by describing it in a sentence, and video moment search, which locates the point in a video where a specific scene appears. Separately, on Mac there's a meeting assistant called Foresight that takes notes entirely on-device and lets users search across both the written transcript and the original recording.

For developers who want to build their own tools, the weights can be downloaded from both Hugging Face and Kaggle. A number of runtimes already support the model out of the box, including LiteRT, MediaPipe, transformers.js, llama.cpp and Ollama, and it can also run directly in a browser via WebGPU. An Android ML Kit interface is coming within the next few weeks, and once it ships it will be able to tap the phone's NPU for hardware acceleration.

For developers in China

With Apache 2.0 licensing, plus existing support from tools like Ollama and llama.cpp, teams in China can pull the model directly and run it in a private deployment without waiting on anyone's approval. That's particularly useful for data-sensitive scenarios such as searching a photo gallery, searching meeting recordings, or building an internal enterprise knowledge base where data isn't supposed to leave the company network. Direct access to Hugging Face from within China tends to be unstable, though, so downloads there usually have to go through a mirror site instead.

China already has no shortage of comparable open-source models of its own — a multimodal embedding model previously open-sourced by Tencent's WeChat team, for instance, once topped the MMEB-v2 leaderboard. The real difference here is scale: EmbeddingGemma 2 was designed from the very beginning around the memory limits of a phone, and that 567MB figure is squarely aimed at on-device use rather than server deployment. As for Chinese-language performance specifically, Google only says its multilingual capability is on par with the first generation and doesn't give a separate score for Chinese-language retrieval — so it's worth running a test against your own corpus before adopting it in production.

Sources: Google's official blog, CocoLoop, Google Developers Blog; parameter composition, context length and the MTEB Code score verified against the Google DeepMind model page, on-device memory and image-processing speed verified against the developer blog.