EmbeddingGemma 2 ships: 740M open multimodal embedder, Apache 2.0

Five modalities in one vector space, runs on a phone at ~190MB RAM, 6x cheaper vector storage via Matryoshka, and a real jump on code retrieval.

Nowline OCT 7 2:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • One small model, five modalities

    Google DeepMind's EmbeddingGemma 2 maps text, code, images, video, and audio into a single 768-dim space at just 740M params, and Google claims it beats embedders twice its size (61.36 on MTEB multilingual v2). Apache 2.0 weights are on Hugging Face and Kaggle, so you can embed once and ship it inside a product with no per-call API bill.

  • Runs on a phone, cuts your vector bill

    Load only the encoders you need, from 270M text-only (~191MB RAM on a Pixel) up to 567MB for the full stack, then truncate output from 768 to 512/256/128 dims via Matryoshka for up to 6x less vector storage. It serves through Ollama, llama.cpp, vLLM, sentence-transformers, and LiteRT, making genuinely offline on-device RAG and semantic search practical this weekend.

  • Code retrieval got a real jump

    EmbeddingGemma 2 scores 78.68 on MTEB Code v1, up 9.92 points over v1, with an 8,192-token context window (4x the old 2,048). For builders that means embedding whole files or large chunks of a repo for code search and RAG using one tiny open model instead of a frontier embedding API.

  • One index for every modality

    A single context can hold up to ~29 images, ~58 video frames, or ~5.5 minutes of audio alongside text, all in the same vector space, across 100+ languages. You can build search that spans docs, screenshots, and recordings without standing up a separate embedding model and index for each one.