Google has released EmbeddingGemma 2, an open-weight multimodal embedding model with 740 million parameters under the Apache 2.0 license, allowing commercial use. The model processes text, code, images, video, and audio within a shared 768-dimensional vector space, featuring an 8,000-token context window that is four times larger than the original version. The text-only configuration runs on devices like Pixel with approximately 191MB of RAM, while the MTEB Code benchmark shows a 9.92-point improvement. The MRL (Matryoshka Representation Learning) feature can reduce dimensions by up to six times to lower storage requirements and speed up lookups. The model weights are available on Hugging Face and Kaggle.
Source: Read the original article

