Gemini Embedding 2: SOTA multimodal embedding model (Google product announcement)
A product-launch tweet from Logan Kilpatrick (Google AI Studio) announcing Gemini Embedding 2, a new SOTA multimodal embedding model that maps text, images, video, audio, and documents into a single shared embedding space. Closed, API-only via Google AI Studio / the Gemini API; no technical report, no benchmark numbers in the tweet, no weights. Filed as a tracking pointer for Google’s closed multimodal embedding stack — the closed-source counterpart to open releases like Qwen3-VL-Embedding.
Key claims
Section titled “Key claims”- Gemini Embedding 2 is positioned as Google’s new SOTA multimodal embedding model [tweet body].
- The model embeds text, images, video, audio, and documents into the same embedding space [tweet body].
- Release is via Google’s API surface (consistent with Google AI Studio / Gemini API delivery for the Gemini family); no weights mentioned [tweet body].
Method
Section titled “Method”Not disclosed. The tweet has no architectural, training, or evaluation detail — only the product name, the modality list, and a single sample image. Mapping five modalities (text, image, video, audio, document) into one embedding space implies modality-specific encoders sharing a contrastive target, but the tweet doesn’t say. Lineage is presumably the Gemini Embedding family previously documented as a closed text-embedding API; the “2” version adds (or makes prominent) the non-text modalities.
Results
Section titled “Results”None reported. No benchmark table, no comparison against MTEB / MMEB / video-text retrieval benchmarks, no ablation. The 861.6K view count is the only quantitative signal on the page.
Why it’s interesting
Section titled “Why it’s interesting”Gemini Embedding 2 is the closed-source multimodal-embedding frontier baseline that open releases this wiki tracks position against — most directly Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking (the open Qwen3-VL-Embedding + Reranker pair, 2B/8B, with full tech report and multi-backend serving), and the smaller HF model-card forms Qwen3-VL-Embedding-8B and Qwen3-VL-Reranker-8B. Filing the announcement gives future open-release pages a stable pointer for what “Gemini Embedding” actually was at the time of comparison. The Open foundation-model releases concept page already names “Gemini Embedding” in its closed-but-API-accessible cohort as a comparison baseline — this page is the artifact behind that mention. Notable also as a Logan Kilpatrick announcement in the same window as Nano Banana 2 / Gemini 3.1 Flash Image (Google product announcement) (Nano Banana 2 / Gemini 3.1 Flash Image), suggesting Google is shipping the Gemini-3.x family one capability at a time via @OfficialLoganK product tweets.
See also
Section titled “See also”- Open foundation-model releases — Gemini Embedding 2 is the closed multimodal-embedding baseline open releases position against
- Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking — open counterpart: Qwen3-VL-Embedding + Reranker tech report (text + image + video, no audio yet)
- Qwen3-VL-Embedding-8B — open Qwen3-VL-Embedding-8B model card
- Qwen3-VL-Reranker-8B — paired open reranker for recall→rerank pipelines
- Nano Banana 2 / Gemini 3.1 Flash Image (Google product announcement) — same author, same Gemini-3.x rollout cadence (Nano Banana 2 / Gemini 3.1 Flash Image)
- Gemini 3 Deep Think: Advancing science, research and engineering — same Gemini 3.x family, reasoning side