dArtBook a call
all news
TechCrunch · May 19, 2026

Google’s Omni Model Blends Text, Audio, and Images Into Video—With a Consumer-First Twist

TechCrunch
Google’s Omni Model Blends Text, Audio, and Images Into Video—With a Consumer-First Twist
May 19, 2026

Google took the wraps off Gemini Omni at its I/O conference, a new family of multimodal models that can take any combination of text, images, audio, and video and produce a coherent video output. The system doesn’t just stitch media together; it reasons across inputs to generate clips that demonstrate an understanding of physics, culture, and history. For example, given the prompt “a claymation explainer of protein folding,” Omni produced a stop-motion video with a voice-over describing amino acid chains and beta sheets.

The first model in the family, Gemini Omni Flash, rolls out today to the Gemini app, YouTube Shorts, and Google’s creative studio Flow. It can generate up to 10 seconds of video—a deliberate limit to encourage broad consumer use, according to DeepMind’s Nicole Brichtova. Longer clips are planned. Users can also create digital avatars by recording themselves speaking a series of numbers, then use those avatars in personalized videos, such as accepting an award or visiting the moon. All outputs include Google’s SynthID watermark to verify authenticity.

Google is positioning Omni Flash primarily as a consumer tool, but enterprise implications are clear. An API will be available in coming weeks, and the model’s ability to render accurate text in video—useful for slogans and product shots—makes it attractive for advertisers and filmmakers. A more capable Omni Pro model is expected later, though Google hasn’t set a release date. For now, the company is betting that ease of use, not raw power, will drive adoption.

Source: TechCrunch

Want a self-updating feed like this on your site?

dArt Studio installs AI for local businesses in Broward & Palm Beach County, FL. We reply within 1 business hour.