Gemini Omni 1.1 Flash: 40-second scenes with first/last-frame control
Google's video model reads 10s of context, upscales to 4K, and drafts at a third the cost. Plus: IBM's Granite 4.2 reasoning models go local, Apache 2.0.

Copy markdown
Scenes now stretch to 40 seconds
Extend a clip in 10-second steps up to a 40-second total, with the model reading up to 10 seconds of prior context instead of just the last frame. Motion, lighting, and characters hold together shot to shot.
You set the first and last frame
Hand Omni a start frame and an end frame and it fills the motion between them, so camera orbits, dolly-zooms, and seamless loops become deterministic instead of prompt-roulette.
Draft at 360p, then upscale
360p drafts render up to 60% faster at a third the cost (~$0.03/s). Lock the shot, then upscale to 720p, 1080p, or 4K. 720p runs about $0.10 a second ($17.50 per 1M video tokens, $1.50 input).
Callable from the API today
Live now in the Gemini API, Google AI Studio, and the Enterprise Agent Platform, plus Google Flow for AI Plus/Pro/Ultra. Adobe, Figma, and Runway are already shipping on it.
One catch: no audio
Voice editing and audio references are unsupported, and any sound inside a reference clip is ignored. Plan to add music and dialogue in post rather than expecting Omni to score the cut.
Elsewhere: IBM Granite 4.2 goes local
IBM shipped Granite 4.2 open reasoning models at 3B/8B/30B under Apache 2.0, with native step-by-step thinking and agentic RL tuned for software-engineering and terminal tasks. The 8B just landed on OpenRouter and runs on Ollama or LM Studio.