Google's new Gemini video model is built for developers, not the feed — Omni 1.1 Flash is generative video growing up
The 27 August release adds the unglamorous things that make AI video usable: 40-second coherent clips, keyframe control, 4K and a cheap draft mode, all GA on the Gemini API. Plus Gemini 3.5 Transcribe.

The viral phase of AI video is winding down. The interesting phase — where it becomes something developers can actually build on — is what Google shipped this week, and it looks a lot less exciting than a talking-avatar demo. That is rather the point.
On 27 August, Google made Gemini Omni 1.1 Flash generally available on the Gemini API, a developer-facing update to its video-generation model. There is no consumer fanfare here. Instead there is a list of the unglamorous features that separate a party trick from a production tool.
What actually changed
The headline additions are about control and coherence, the two things generative video has always been worst at:
- Longer, coherent clips. The model now reads up to 10 seconds of prior context when extending a video — up from a single second — so scenes can be stretched in 10-second increments to around 40 seconds while holding visual consistency and narrative thread. Length without drift is the hard part, and this is a direct attack on it.
- Keyframe control. First-and-last-frame interpolation lets a developer specify a starting and ending image and have the model generate the continuous motion between them. That is the difference between "generate something" and "generate exactly this transition."
- 4K output alongside 1080p, aimed at professional production rather than social clips.
- A cheap draft mode. A new 360p preview tier runs up to 60% faster and at roughly a third of the cost of the standard 720p output — so teams can iterate cheaply and only spend on the final render.
- Video references, letting a prompt point at up to a few seconds of existing footage to keep a character or style consistent across shots.
The transcription half
It did not arrive alone. The day before, on 26 August, Google took Gemini 3.5 Transcribe to general availability — a speech-to-text model with utterance-based language detection across more than 85 languages, speaker diarization, word-level timestamps and custom vocabulary biasing. Unshowy, and exactly the kind of building block that ends up quietly inside a lot of other products.
Why the boring release is the important one
None of this will trend the way a May-style "look what it generated" launch does. But the shift is real: the features Google chose to ship — draft modes, keyframe control, resolution tiers, longer coherent shots — are the features a studio or an app developer asks for, not the ones that make a good demo reel.
That is how a capability stops being a novelty and becomes infrastructure. AI video spent a year proving it could make something eye-catching in one shot. The 1.1 releases are about making it controllable, affordable and long enough to use — which is the far less viral, far more consequential problem to solve.
Ask Relay — he reads every question himself and replies personally by email.
