PhysStream turns one image into video you steer as physics plays out

Penn, Snap and KAUST's model takes velocity nudges to redirect objects mid-shot, cuts trajectory error 12% and wins 85% of head-to-heads. Code's still pending.

Nowline SEP 16 12:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • One image in, a video you can push around

    PhysStream is autoregressive image-to-video: feed a single frame, then steer objects mid-generation with sparse “velocity-increment” signals — nudge a ball and the scene keeps its physics consistent. A structured scene memory (position and object-tracking maps) is what holds multi-object rigid-body dynamics together frame to frame.

  • The numbers: 12% straighter trajectories, 85% preferred

    On synthetic benchmarks it cuts trajectory error ~12% and motion-distribution distance (FVMD) ~33% against the strongest baselines, and human raters picked it in over 85% of in-the-wild comparisons. It's from UPenn, Snap and KAUST and is accepted to SIGGRAPH Asia 2026.

  • The catch: code is 'coming soon,' scope is tabletop

    You can't run this today — the project page lists code and datasets as “coming soon,” and the demos are rigid-body tabletop scenes, not open-world footage. Treat it as a signal of where controllable video is heading, not a tool to ship with yet.

  • Why it matters when the weights land

    Interactive, physics-aware control is the missing piece for using generated video anywhere deterministic — product demos, game-feel prototyping, ad mockups where the note is “make the bottle roll left.” Drag-to-direct turns that from a render-farm job into a prompt-and-nudge workflow.