
EmbeddingGemma 2 brings multimodal embeddings to local devices
Google DeepMind has released EmbeddingGemma 2, a lightweight model designed to run locally on consumer hardware. It natively maps text, images, audio, and video into a shared space to power offline search and retrieval.
Published by Jin · 2 min read · 7 OCT 2026
- 308M
- roughly 100M
- 200M
- 768-dimension (truncatable to 128, 256, or 512)
- sub-200MB

Google DeepMind has released EmbeddingGemma 2, a compact model designed to run directly on local consumer hardware. The model natively maps combinations of text, images, audio, and video into a unified vector space, allowing applications to search and connect information without sending data to the cloud.
Multimodal capabilities on edge devices
Built on the Gemma 4 architecture, EmbeddingGemma 2 contains 740 million total parameters and is released under an Apache 2.0 license. It expands significantly upon its predecessor by supporting code, images, video, and audio alongside text. Developers can use it to find specific video clips using voice memos or search audio recordings with text queries.
The model is modular by design. It requires 270 million parameters for text-only workloads, with optional vision (170 million) and audio (300 million) encoders for full multimodal support. Using Matryoshka Representation Learning — a technique that allows vectors to be safely truncated without losing core meaning — developers can dynamically reduce output dimensions from 768 down to 512, 256, or 128. This provides up to a sixfold reduction in storage space for local vector databases.
Performance and deployment
EmbeddingGemma 2 supports an 8K token context window, which is four times larger than the first version. This allows the system to process up to 5.5 minutes of audio, 29 images, or 58 video frames locally. On a Google Pixel 11 Pro, the model requires roughly 191MB of active RAM for text-only weights and about 567MB for the full multimodal setup when quantized.
The model achieves leading benchmark scores among sub-1B multimodal embedders, including strong performance on MTEB Code and the Massive Audio Embedding Benchmark. It enables fully offline retrieval-augmented generation pipelines — systems that feed local data to language models to improve accuracy — when paired with generative models like Gemma 4. Developers can access the model weights on Hugging Face and Kaggle.
Source — Original announcement ↗
Worth a read?
Comments · 0