EmbeddingGemma 2 Brings Private Multimodal Search to Phones and Laptops
Google's EmbeddingGemma 2 is a 740M-parameter open model that runs multimodal search for text, images, audio and video entirely on-device.
Google DeepMind has released EmbeddingGemma 2, a 740-million-parameter open model that turns a phone or laptop into a search engine for its own files, with no server in the loop. The launch came on October 6 under an Apache 2.0 licence, and it is the first entry in Google's small EmbeddingGemma line to handle anything beyond text. Google built it on the Gemma 4 architecture and calls it the most capable model available for on-device multimodal embeddings.
The practical change is smaller than the framing but more useful. A voice memo can find the matching moment in a video. A typed sentence can surface a frame or a photo. Hours of recorded audio become searchable by description. All of it runs locally, which is the point: EmbeddingGemma 2 targets privacy-first retrieval-augmented generation and semantic search rather than chat.
Inside EmbeddingGemma 2's 740 million parameters
The headline number hides a modular design. Only 270 million parameters belong to the text backbone; a 170-million vision encoder and a 300-million audio encoder sit alongside it, and an app can load only the pieces it needs. Everything the model sees, whether a line of Python, a video frame or a spoken sentence, is compressed into the same 768-dimensional vector space, so a text query can be compared directly against an image or an audio clip.
| Specification | EmbeddingGemma 2 |
|---|---|
| Total parameters | 740M (270M text, 170M vision, 300M audio) |
| Embedding dimensions | 768, truncatable to 512, 256 or 128 |
| Context window | 8K tokens, four times the previous version |
| Active RAM | About 191MB text-only, about 567MB full multimodal |
| Architecture | 24 layers, hidden size 2048, 262,144-token vocabulary |
| Licence | Apache 2.0 |
Two design choices carry most of the engineering weight. The context window has grown to 8,000 tokens, four times the text-only first-generation model, which is what makes long transcripts and multi-image inputs workable. The model also uses Matryoshka Representation Learning, a training method that lets developers cut a vector from 768 dimensions down to 512, 256 or 128 without retraining. Smaller vectors mean less storage per indexed item, which matters when a library of 50,000 photos has to fit inside a phone's memory budget.
The generational jump is easy to understate. The first EmbeddingGemma, released in 2025, handled text alone. This version folds code, images, video and audio into one space and quadruples the context window to 8,000 tokens. The two models differ in modality coverage far more than in size.
Why the memory numbers matter more than the parameter count
Google's own figures put text-only weights at roughly 191MB of active RAM once quantized, rising to about 567MB for the full multimodal stack on a Pixel 11 Pro. That is the number developers will plan around. Vision and audio encoders are the expensive part, and the modular layout means a note-taking app can skip both.
A second saving appears only when EmbeddingGemma 2 is paired with Gemma 4. The two share a text tokenizer and an audio encoder, so an on-device retrieval-augmented generation setup running both needs less memory than two unrelated models would. For anyone building a local assistant that answers questions about a user's own files, that overlap decides whether the stack fits on a flagship phone.
App design follows from the vector space. Once a photo, a spoken note and a snippet of code share one coordinate system, a single index can serve many query types, and developers stop building a separate pipeline for each modality. That reduces both code and the number of models a phone has to keep resident.
The trade-offs developers will hit
Several constraints are worth reading before the launch enthusiasm sets in. Inference requires bfloat16 or float32 precision, and float16 is not recommended, which rules out some older mobile accelerators. Input limits are firm rather than generous: roughly 29 to 114 images, about 58 video frames or around 327 seconds of audio per pass, so long recordings have to be chunked and indexed in pieces.
Training data runs to January 2025, so anything newer sits outside the model's knowledge unless it arrives as an input. The economics cut both ways. A local model removes per-query API billing and network round trips, and pushes model size, updates and storage onto the device instead.
Quality claims are harder to verify than the memory claims. Google's materials frame the gains in terms of footprint, latency and RAM, while public head-to-head benchmarks against larger hosted embedding models are not the centrepiece of the release. For a developer choosing between a cloud embedding API and a local model, that gap is the real decision point.
For users, the visible effect is narrower than the marketing implies. On-device embeddings remove the need to upload a query or an index, which is not the same as making an app private. Whether a local index stays private depends on how the app stores and syncs the data it keeps on the device, and the Apache 2.0 licence governs none of that.
What is available now
Tooling is not a bottleneck. Weights are on Hugging Face and Kaggle, with quantized GGUF builds from Unsloth, and the model runs through sentence-transformers, PyTorch, JAX, Keras, llama.cpp and Ollama. Google is also shipping it into Google AI Edge MediaPipe and LiteRT, with Vertex AI Model Garden to follow. Two features built on the model are coming to Google AI Edge Gallery, the company's app for testing models on a device.
The model card lists more than 100 supported languages and reports strong gains on code tasks. Google DeepMind research engineers Sahil Dua and Henrique Schechter Vera detailed the release.
Why this matters
On-device AI has spent two years being measured by what a chatbot can say. EmbeddingGemma 2 points at a different question: what a device can find. If retrieval moves onto the phone, privacy becomes a property of the hardware rather than a policy promise, and app design shifts from sending queries to a server toward indexing what is already on the device. The 567MB footprint is the constraint that decides which of those apps ship this year.
Sources
EmbeddingGemma 2: an open, lightweight multimodal embedding model
EmbeddingGemma 2 model card | Google AI for Developers
google/embeddinggemma-2 - Hugging Face
unsloth/embeddinggemma-2-GGUF - Hugging Face
README.md · unsloth/embeddinggemma-2-GGUF at main
EmbeddingGemma 2: an open, lightweight multimodal embedding model
Related Articles
- Google DeepMind Brings Gemma 4 Open Models to Amazon Bedrock for Enterprise AI
- Google and NVIDIA Launch DiffusionGemma to Deliver 4x Faster Parallel Text Generation
- Google Expands Gemini API File Search with Multimodal Support and Citations
✔Human Verified
Researched and cross-referenced against primary sources by the Bytevyte editorial team. This article was generated with the assistance of artificial intelligence and reviewed by the Bytevyte editorial team.