Grok Imagine 1.5 adds 1080p text-to-video and character references
Full-HD clips with native audio, up to seven references to lock a face and voice across scenes, callable on the xAI API — plus OpenAI adds a Sol Fast mode.

Copy markdown
1080p video straight from a text prompt
Grok Imagine Video 1.5 now generates native 1080p from text alone — no starting image — as well as from images, across web, iOS, and Android, with sound effects, ambience, and dialogue produced in the same pass. Announced Aug 1.
Seven references keep a character consistent
You can anchor a generation with up to seven image references — character, product, or location — and pair a character image with a voice sample to hold the same face and voice across scenes. That is the missing piece for stitching multi-shot stories that don't fall apart between cuts.
It's on the API: grok-imagine-video-1.5
The xAI API exposes text-to-video, image references, and 1080p through the grok-imagine-video-1.5 model; voice references take a separate request. Enough to wire a shot-list-to-video pipeline this weekend without touching the app.
The catch: gated to SuperGrok first
Image and voice features launched for SuperGrok Heavy and Plus subscribers in the US, with rollout to all tiers promised over the next few days. If you're on a lower tier, the API is your fastest way in right now.
Elsewhere: OpenAI's Sol gets a Fast mode
GPT-5.6 Sol now offers a Fast mode — up to 2.5x faster at twice the price — that replaces Priority Processing, and existing priority-tagged requests keep working. Reach for it when Sol latency, not cost, is your bottleneck.