LTX-2.5 ships open weights: local AI video with synchronized audio
The 22B model runs in ComfyUI day one, quantized to fit consumer GPUs — just read the $10M license line first. Plus: on-device sign-to-text on Pixel 11.

Copy markdown
Open weights, and the audio comes with it
Lightricks' LTX-2.5 is a 22B open-weights model that generates synchronized video AND audio from text, image, or video prompts — clips of roughly 2 to 20 seconds, up to 4K at 50fps, with native multi-shot scenes that hold character and style across cuts. No API bill and no per-second metering: the weights are on Hugging Face.
6.8 seconds on a superchip — or a GGUF on your card
The distilled model renders a 10-second clip from an image in about 6.8 seconds on NVIDIA superchips. For everyone else, INT8/fp8/NVFP4/GGUF builds plus ComfyUI day-zero templates bring it down to consumer GPUs — a local text-to-video-with-sound rig you can stand up this weekend.
Read the $10M license line before you ship
LTX-2.5 uses the LTX-2 Community License: free for commercial use if your organization is under $10M in annual revenue, paid licensing above that. Check it before you wire it into a product you plan to sell.
Elsewhere: sign-to-text goes on-device
Google DeepMind shipped SL2T, a sign-language-to-text model, into Gboard and Live Transcribe on Pixel 11 — translating ASL to English from on-device pose landmarks, so no video leaves the phone. There's no public API or open weights yet, but it's the first time this lands in a shipping consumer keyboard.