Google launched EmbeddingGemma 2, a lightweight multimodal model capable of handling multiple tasks while outperforming larger models. The predecessor contained 308 million parameters and ran on less than 200MB of RAM thanks to quantization-aware training. The model topped MTEB leaderboards in Multilingual v2, English v2, and Code among open models under 500 million parameters. It supported over 100 languages and offered a 2,000-token context window. Inference ran at 15 to 22 milliseconds on EdgeTPU hardware, making it suitable for privacy-sensitive offline applications.
Source: Read the original article

