LTX-2.5: open-weights 4K video with audio, runs on a 16GB GPU
Lightricks' world model makes 4K clips with synced audio in seconds, ships day-one in ComfyUI, and is free under $10M ARR — plus a robotics checkpoint.

Copy markdown
One image in, a clip with sound out
LTX-2.5 turns a prompt or a single still into video with synchronized audio in one pass — text-to-video, image-to-video, video-to-video and audio-to-video all in the same model. Lightricks put the open weights on Hugging Face on Aug 11.
Runs on your machine, ComfyUI on day one
The ~22B model ships as a base plus an NVIDIA-tuned distilled build, running from datacenter GPUs down to a 16GB card or a Mac, with native ComfyUI integration and quantized checkpoints from the start. No API round-trip needed to generate.
Fast enough to iterate on shots
A 10-second 720p image-to-video clip renders in ~6.8s on NVIDIA GB200 superchips, or ~23.7s at 1080p on LTX's managed API. That's fast enough to tweak a shot and re-roll instead of queuing renders overnight.
Native 4K HDR, audio baked in
It outputs up to 4K HDR and synthesizes video and audio jointly with modality-aware guidance, holding characters, scenes and voices consistent across multi-shot cuts — a short-film pipeline in a single weight file.
A world model you can fine-tune for robotics
Beyond clips, LTX-2.5 ships a pretrained physical-AI checkpoint you can fine-tune to generate synthetic training data and simulate environments; Markov Robotics is already using it. Free for orgs under $10M ARR — larger teams license.