Gemini Embedding 2: First Natively Multimodal Embedding Model Unifies Text, Image, Video, Audio into Single 3072-D Space
Google Blog·high signal
Google launched Gemini Embedding 2 on March 10, 2026 — the first embedding model that natively ingests text (8,192 tokens), images (up to 6 per request), video (120s), and audio without transcription intermediaries, producing a single 3,072-dimensional vector. It scores 68.32 on MTEB English, a 5.09-point margin over previous leaders, and 68.8 on video retrieval benchmarks. For RAG pipelines handling mixed-media corpora, this eliminates separate embedding pipelines per modality.