Google unveiled Gemini Omni at its I/O developer conference in May 2026, a new family of multimodal AI models designed to generate video, images, and audio from any combination of inputs. CEO Sundar Pichai described the system as able to “create anything from any input.”
Unlike Google’s existing video model Veo, which converts text and images into video, Omni reasons across all input types simultaneously to produce a consistent output. Nicole Brichtova, director of product management at Google DeepMind, called it “the next step towards the progression of combining the intelligence of Gemini with the rendering capabilities of our media models.” The long-term vision includes generating images from audio and audio from video.
The first model in the family, Gemini Omni Flash, rolled out on the day of the announcement to the Gemini app, YouTube Shorts, and AI creative studio Flow. Flash renders up to 10 seconds of video — a deliberate product decision rather than a technical limit, according to Brichtova, based on anticipated consumer usage patterns. A more capable Omni Pro model is planned but has no confirmed release date.
Omni also supports user-generated digital avatars, a feature available on YouTube Shorts at launch. To reduce deepfake risk, users must complete an onboarding process that involves recording themselves and speaking a series of numbers. All videos generated with Omni will carry Google’s SynthID digital watermark. Brichtova noted that editing prompts need to be highly specific, as vague instructions risk unintended alterations to video elements.
Google will make Omni available via API in the coming weeks. Brichtova pointed to advertising and filmmaking as areas where the model’s capabilities — including accurate text rendering — could see professional adoption. Startup Luma AI is building a comparable tool powered by its own unified model.
Google’s broader rationale, as stated by Pichai, is that training AI on multiple formats gives it “a deeper understanding of the world,” and that Omni represents a step toward models that simulate reality rather than predict text.
Source: TechCrunch