EmbeddingGemma 2 is an open multimodal embedding model from Google DeepMind that maps text (including code), images, video, and audio inputs into a unified 768‑dimensional vector space. It comprises a 270 M‑parameter text encoder, a 170 M‑parameter vision encoder, and a 300 M‑parameter audio encoder, totaling 740 M parameters. Designed for efficiency, it can run on consumer hardware such as mobile devices and laptops.

Read original