Skip to content

Introducing Gemini Omni

Google’s I/O 2026 announcement video for Gemini Omni, a new natively multimodal family that takes text + image + audio + video as input and (so far) produces video as output. The first ship is Gemini Omni Flash, framed as the video-side counterpart to Nano Banana: Gemini’s reasoning stack fused with a generative-media head in a single forward pass, with conversational multi-turn editing as the headline interaction pattern. Available in the Gemini app, Google Flow, YouTube Shorts and YouTube Create from launch; developer API “in the coming weeks” via AI Studio / Vertex AI. The companion model card describes the model as a transformer (Vaswani et al. 2017) with native multimodal support for text/vision/video/audio inputs; no parameter count, training corpus, or benchmark numbers are released.

  • Gemini Omni is positioned as a unified “create anything from any input” model family, with the first model (Omni Flash) accepting text + image + audio + video and emitting high-resolution video with audio; image and audio output are stated as future modalities [video / official blog].
  • Architecturally described as a transformer with native multimodal support for text, vision, video and audio inputs — Gemini’s reasoning stack and the generative-media head live in the same forward pass rather than a relay of specialized systems [Gemini Omni Flash model card].
  • Headline interaction pattern is conversational multi-turn video editing — each instruction builds on the last, with claimed character / scene / physics consistency across turns [video / official blog].
  • Omni Flash is rate-limited at launch to ~10-second clips, framed as a product decision (distribution to consumer surfaces first) rather than an architectural ceiling; longer durations stated as “in the pipeline” [TechCrunch coverage / official blog].
  • Every Omni-generated video carries a SynthID watermark, verifiable via Gemini app / Gemini in Chrome / Google Search [video / official blog].
  • Distribution at launch: Gemini app + Google Flow for Google AI Plus / Pro / Ultra subscribers; free in YouTube Shorts and YouTube Create; developer + enterprise APIs in “the coming weeks” via AI Studio and Vertex AI [video / official blog].
  • No benchmark numbers (T2V, I2V, R2V/A, video editing, image generation) are released; the model card explicitly defers all evaluation tables to the API rollout [Gemini Omni Flash model card].
  • Model card names remaining limitations: maintaining complete consistency across edits, generating scenes with complex motion, and rendering perfectly accurate text [Gemini Omni Flash model card].

The artifact is an announcement video — no transcript was available at fetch time, and no architecture paper or model report is published. The companion model card describes Omni Flash as “a transformer-based model (Vaswani et al., 2017) with native multimodal support for text, vision, video and audio inputs” producing “high-quality, high-resolution video with audio”. Beyond that one sentence, all technical structure (tokenizers, vision/audio interface, training objective, post-training stack, parameter count, training corpus) is undisclosed. Google DeepMind’s framing positions the model as a “world model” whose generative rollout is governed by Gemini’s reasoning trace plus an intuitive grasp of physics (gravity, kinetic energy, fluid dynamics) — distinguishing it rhetorically from frame-extrapolation video generators like Veo. Treat all of this as positioning, not measurement, until the API + eval tables ship.

None disclosed quantitatively. The video and supporting blogs ship qualitative demos only — claymation explainer of protein folding, hippocampus stop-motion, mirror-ripple person-to-reflective-material, surreal background swaps, “Avatars” digital-self video. No comparison numbers against Veo 3.1, Sora-2, Wan 2.7, Hailuo 2.3, or any other frontier video model. No human-eval study. The Omni Flash model card commits to publishing T2VA / I2VA / R2VA / video editing / image generation evaluations at the developer-API rollout, not before.

This is the closed-flagship counterpart to OmniWeaving (OmniWeaving: Towards Unified Video Generation with Free-form Composition and Reasoning) — the first filed unified-multimodal model on the wiki where the output modality is video — landed as an announcement rather than a paper. The TestingCatalog leak from May 2 was Gemini 'Omni' leak: new omni model spotted on the video-generation tab — TestingCatalog; this is the actual reveal. For Unified Multimodal Models it’s the first frontier closed model whose declared output modality is video (with image/audio output stated as roadmap), and for World Foundation Models it slots in alongside Project Genie: Experimenting with infinite, interactive worlds as another closed-flagship DeepMind world-model product — but with a generative-rollout video output surface rather than Genie 3’s interactive-rollout web prototype. Compared to Qwen3.5-Omni (Qwen3.5-Omni Technical Report) and Ming-flash-omni 2.0 (Ming-flash-omni 2.0) — open omni models that include audio/text generation but not high-fidelity video — Omni Flash is the API-accessible side of the same architectural bet, but with the opposite trade: video output is in, but weights, training data, parameter counts, and benchmarks are all out.