EmbeddingGemma 2 is an open‑source multimodal embedding model released under the Apache 2.0 license, built on the Gemma 4 architecture and containing 740 million parameters. It supports text‑only workloads with a core of 270 M parameters, while optional vision (170 M) and audio (300 M) encoders can be added for full multimodal capability. Using Matryoshka Representation Learning, output vectors can be truncated from 768 dimensions to 512, 256, or 128 dimensions, delivering up to a six‑fold reduction in storage and memory for local vector databases. Quantized inference on a Google Pixel 11 Pro consumes roughly 191 MB of RAM for text‑only weights and about 567 MB for the complete multimodal model. The model features an 8K‑token context window, enabling it to process up to 5.5 minutes of audio, 29 images, 58 video frames, or equivalent interleaved combinations directly on edge hardware. Benchmark results show a 9.92‑point gain on MTEB Code (rising from 68.76 to 78.68) and leading performance among sub‑1B multimodal embedders on MTEB and MAEB, often matching or exceeding much larger models. EmbeddingGemma 2 retains the multilingual text strength of its predecessor and facilitates on‑device retrieval‑augmented generation when paired with Gemma 4, ensuring privacy‑preserving, low‑latency cross‑modal search. Model weights are available on Hugging Face and Kaggle, with deployment options via MediaPipe, LiteRT, transformers.js, WebGPU, and various inference servers, and fine‑tuning guidance is provided by Unsloth.

Read original