Skip to content

Gemini Embedding 2: SOTA multimodal embedding model (Google product announcement)

A product-launch tweet from Logan Kilpatrick (Google AI Studio) announcing Gemini Embedding 2, a new SOTA multimodal embedding model that maps text, images, video, audio, and documents into a single shared embedding space. Closed, API-only via Google AI Studio / the Gemini API; no technical report, no benchmark numbers in the tweet, no weights. Filed as a tracking pointer for Google’s closed multimodal embedding stack — the closed-source counterpart to open releases like Qwen3-VL-Embedding.

  • Gemini Embedding 2 is positioned as Google’s new SOTA multimodal embedding model [tweet body].
  • The model embeds text, images, video, audio, and documents into the same embedding space [tweet body].
  • Release is via Google’s API surface (consistent with Google AI Studio / Gemini API delivery for the Gemini family); no weights mentioned [tweet body].

Not disclosed. The tweet has no architectural, training, or evaluation detail — only the product name, the modality list, and a single sample image. Mapping five modalities (text, image, video, audio, document) into one embedding space implies modality-specific encoders sharing a contrastive target, but the tweet doesn’t say. Lineage is presumably the Gemini Embedding family previously documented as a closed text-embedding API; the “2” version adds (or makes prominent) the non-text modalities.

None reported. No benchmark table, no comparison against MTEB / MMEB / video-text retrieval benchmarks, no ablation. The 861.6K view count is the only quantitative signal on the page.

Gemini Embedding 2 is the closed-source multimodal-embedding frontier baseline that open releases this wiki tracks position against — most directly Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking (the open Qwen3-VL-Embedding + Reranker pair, 2B/8B, with full tech report and multi-backend serving), and the smaller HF model-card forms Qwen3-VL-Embedding-8B and Qwen3-VL-Reranker-8B. Filing the announcement gives future open-release pages a stable pointer for what “Gemini Embedding” actually was at the time of comparison. The Open foundation-model releases concept page already names “Gemini Embedding” in its closed-but-API-accessible cohort as a comparison baseline — this page is the artifact behind that mention. Notable also as a Logan Kilpatrick announcement in the same window as Nano Banana 2 / Gemini 3.1 Flash Image (Google product announcement) (Nano Banana 2 / Gemini 3.1 Flash Image), suggesting Google is shipping the Gemini-3.x family one capability at a time via @OfficialLoganK product tweets.