AI ONLINE6 September 2026
The AI News Desk

RelayON THE WIRE

The whole field of AI — read, checked, and explained.
Models & Releases

Google's new Gemini video model is built for developers, not the feed — Omni 1.1 Flash is generative video growing up

The 27 August release adds the unglamorous things that make AI video usable: 40-second coherent clips, keyframe control, 4K and a cheap draft mode, all GA on the Gemini API. Plus Gemini 3.5 Transcribe.

Priya AnandBy Priya AnandBusiness Editor
28 August 2026
Listen to this postread by Relay

The viral phase of AI video is winding down. The interesting phase — where it becomes something developers can actually build on — is what Google shipped this week, and it looks a lot less exciting than a talking-avatar demo. That is rather the point.

On 27 August, Google made Gemini Omni 1.1 Flash generally available on the Gemini API, a developer-facing update to its video-generation model. There is no consumer fanfare here. Instead there is a list of the unglamorous features that separate a party trick from a production tool.

What actually changed

The headline additions are about control and coherence, the two things generative video has always been worst at:

  • Longer, coherent clips. The model now reads up to 10 seconds of prior context when extending a video — up from a single second — so scenes can be stretched in 10-second increments to around 40 seconds while holding visual consistency and narrative thread. Length without drift is the hard part, and this is a direct attack on it.
  • Keyframe control. First-and-last-frame interpolation lets a developer specify a starting and ending image and have the model generate the continuous motion between them. That is the difference between "generate something" and "generate exactly this transition."
  • 4K output alongside 1080p, aimed at professional production rather than social clips.
  • A cheap draft mode. A new 360p preview tier runs up to 60% faster and at roughly a third of the cost of the standard 720p output — so teams can iterate cheaply and only spend on the final render.
  • Video references, letting a prompt point at up to a few seconds of existing footage to keep a character or style consistent across shots.

The transcription half

It did not arrive alone. The day before, on 26 August, Google took Gemini 3.5 Transcribe to general availability — a speech-to-text model with utterance-based language detection across more than 85 languages, speaker diarization, word-level timestamps and custom vocabulary biasing. Unshowy, and exactly the kind of building block that ends up quietly inside a lot of other products.

Why the boring release is the important one

None of this will trend the way a May-style "look what it generated" launch does. But the shift is real: the features Google chose to ship — draft modes, keyframe control, resolution tiers, longer coherent shots — are the features a studio or an app developer asks for, not the ones that make a good demo reel.

That is how a capability stops being a novelty and becomes infrastructure. AI video spent a year proving it could make something eye-catching in one shot. The 1.1 releases are about making it controllable, affordable and long enough to use — which is the far less viral, far more consequential problem to solve.

Tune your feed
Like to get more stories like this in your For You feed — dislike for fewer.
Sources
Priya Anand — Business Editor. Priya tracks the money and the market: raises, deals, pricing, and the economics shaping where AI goes next. Spot something wrong? Tell me and I'll correct it in public.
Got a question about this?

Ask Relay — he reads every question himself and replies personally by email.

Ask Relay →