Hands-on review

LTX 2.5: open-source 4K video with audio

Lightricks’ LTX-2 line generates native 4K video and synchronized audio together, with open weights you can self-host.

By the Vuela.ai content team ·

Official from Lightricks.

What it nails

  • Native 4K with synchronized audio in one pass
  • Open weights under a permissive license
  • Multishot: connected shots in one generation
  • Clips up to 20 seconds in the Fast variant

Where it struggles

  • Setup and tuning expect technical users
  • Self-hosting 22B at 4K HDR wants serious hardware
  • Prompt fidelity trails the closed premium models
  • Quality depends on your own pipeline and hardware

LTX 2.5 is the latest point release in Lightricks’ open-source LTX-2 video model line. Unlike the closed flagships, it ships with open weights, so studios and developers can self-host, fine-tune, and embed it directly. This review is a specs-and-positioning analysis based on Lightricks’ published documentation and the public open-source release, not a per-prompt lab test.

The headline is that LTX-2 pairs native 4K output with synchronized audio in a single generation. LTX 2.5 moves the line to a 22B model and adds what the earlier releases could not do: multishot scenes that hold their characters across cuts, automatic duration, native 4K HDR, and Diffusion Fidelity Rendering, which spends compute where a scene is complex instead of spreading it evenly across every frame.

What is LTX 2.5?

LTX 2.5 is the August 2026 release of LTX-2, Lightricks’ multimodal open-weights foundation model for video. It is a 22B asymmetric dual-stream diffusion transformer that generates video and audio jointly, targeting native 4K HDR, with clips up to 20 seconds in the speed-optimised Fast variant. The weights are public on Hugging Face, so it runs in ComfyUI, through the LTX API, or in a self-hosted pipeline.

Positioning: LTX-2 is the open-source option to beat when you need 4K plus audio without paying per second or sending footage to a closed API. The trade-off is that you own the setup, the hardware, and the tuning.

How we assess LTX 2.5

This is a capability assessment built from Lightricks’ published specs, the open-source release notes, and how the model is positioned against other 2026 video models. We weigh the dimensions that matter for production use rather than running a single prompt.

  1. Resolution and audio Native 4K with synchronized audio generated in the same pass.
  2. Openness Open weights and tooling that can be self-hosted and fine-tuned.
  3. Scene control Whether multishot and automatic duration hold a scene together across cuts.

The test results

Test 1. Resolution and audio fidelity

LTX 2.5 is documented to produce native 4K with HDR pipelines built in, and to generate audio jointly with the picture through bidirectional cross-attention rather than adding it in a separate pass. On paper this matches what the closed flagships do for resolution and audio sync, which is rare for an open model. The practical ceiling depends on your GPU and the variant you run.

Test 2. Open weights and self-hosting

The full weights and training framework are public, so LTX 2.5 can be embedded in custom tools, fine-tuned on a house style, or run entirely offline. For teams that need data control or per-clip cost control at scale, this is the core reason to pick LTX over a closed API.

Test 3. Multishot and scene length

This is the real jump in 2.5. One generation can cut between shots and keep the same characters and setting across them, and automatic duration lets the model pick the length the description implies. With up to 20 seconds in the Fast variant, a whole idea fits in a single generation instead of being stitched from several.

Where it struggles

Technical setup. Getting the best out of open weights means managing your own environment, models, and hardware.

Hardware. A 22B model at 4K HDR is not a laptop workload: self-hosting the top settings means real GPUs.

Prompt fidelity. The closed premium models still lead on instruction following for the hardest prompts.

Who should use it

LTX 2.5 is the strongest pick when you want 4K plus audio, control over your pipeline, and predictable cost at volume. If you want a no-setup, point-and-shoot experience with the highest prompt fidelity, a closed flagship is the easier path.

How Vuela.ai fits alongside LTX 2.5

You can run LTX 2.5 yourself, or pick it in Vuela.ai with no setup at all: LTX 2.5 Fast is one of the models in the image-to-video and text-to-video tools, native audio included.

Self-host when you need the weights, the data control, or your own fine-tune. Use Vuela.ai when you want the same model plus the full pipeline: clone, translate with lip-sync, add voiceover and ship, with the rest of the top models on one plan.

Open-source power, production-ready pipeline

Vuela.ai bundles the best video models with cloner, translator, and 70+ tools on one flat plan.

The verdict

LTX 2.5 is the open-source video model to beat in 2026: native 4K with audio, open weights, and efficiency that keeps it practical on real hardware. The cost is the technical ownership that comes with any self-hosted model.

For open pipelines and cost control, LTX 2.5 is a top pick. For zero-setup convenience, pair it with, or swap it for, a managed platform.

LTX 2.5 review FAQ

Is LTX 2.5 free? +

LTX-2 ships with open weights under a permissive license, so the model itself is free to self-host. You still pay for the hardware or cloud compute you run it on.

Does LTX 2.5 generate audio? +

Yes. LTX-2 generates synchronized audio together with the video in a single pass, rather than adding sound in a separate step.

What resolution does LTX 2.5 support? +

It targets native 4K with HDR built in. Clip length reaches 20 seconds in the Fast variant, and the practical ceiling depends on the variant and the hardware you run it on.

Can I run LTX 2.5 myself? +

Yes. The weights and tooling are public, so it runs in ComfyUI, via the LTX API, or in a fully self-hosted pipeline.

Can I use LTX 2.5 inside Vuela.ai? +

Yes. LTX 2.5 Fast is one of the models you can pick in Vuela.ai’s image-to-video and text-to-video tools, with its native audio included, so you get the model plus cloner, translator, and 70+ tools on one plan without self-hosting.

Build your pipeline with Vuela.ai

Flat-rate access to the best models, plus cloner, lip-sync translator, and 70+ tools.