Grok Imagine adds native 1080p text-to-video, voice + face locking

xAI's update puts 1080p on the API at $0.25/sec, pins up to seven scene references, and holds a face and voice steady across shots — best parts still gated.

Nowline AUG 3 11:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • Text-to-video, no seed image needed

    Grok Imagine now synthesizes the opening frame from a text prompt and animates forward — no starting still required. It's live free on grok.com/imagine, iOS, and Android, and the API model string is grok-imagine-video-1.5.

  • Native 1080p — and here's the tab

    Direct HD render, no upscaling, on the autoregressive Aurora architecture. API pricing runs $0.08/sec at 480p, $0.14 at 720p, and $0.25 at 1080p — about $900 for 60 minutes of 1080p a month. Budget before you batch.

  • Lock a face and a voice across every scene

    Feed a character photo plus a voice sample and speaker embeddings keep both consistent shot to shot. That's the piece multi-scene narrative video was missing — your character stops drifting between cuts.

  • Seven references, seven anchors

    Pin up to seven elements — a face, a product, a location, a prop — per generation while everything else varies. That's repeatable enough for product demos and branded shorts, not just one-off clips.

  • The catch: it's gated

    Voice references need a sales inquiry with no self-serve at launch, the free tier is cut out of every new feature, and image and voice refs are rolling out US-first on SuperGrok Heavy and Plus.

  • Elsewhere: voice and app-builders move

    Microsoft previewed MAI-Realtime, a full-duplex voice model that can run web search and tools mid-conversation (preview, no GA date). And Google scrapped its standalone AI Studio app after ~800K pre-orders, folding app-building straight into Gemini.