LTX-2.5 - LTX, an open-source AI video generation foundation model.
LTX-2.5 is LTX's open-source AI video generation foundation model, featuring 22 parameters with open weights upon release. The model supports native multi-camera generation, 4K HDR and RAW/EXR professional workflows, and can generate audio simultaneously. ...
What is LTX-2.5?
LTX-2.5 is LTX's open-source AI video generation foundation model, featuring 22 parameters with open weights upon release. The model supports native multi-camera generation, 4K HDR, and professional RAW/EXR workflows, can generate audio simultaneously, and possesses precise video editing capabilities. At a time when closed-source models are becoming increasingly expensive, LTX-2.5 offers a professional-grade open-source alternative to the AI video field through its open ecosystem strategy.
Main functions of LTX-2.5
- Native multi-camera generationIt can generate multiple consecutive shots with different camera angles and perspectives in a single generation, maintaining a high degree of consistency in character appearance, environmental atmosphere, lighting conditions and sound between shot transitions, thus solving the core pain point of shot continuity in AI story creation.
- Diffusion Fidelity RenderingIt adaptively allocates rendering computing power based on scene complexity, giving more computing resources to complex and dynamic shots, while automatically saving resources for static or simple scenes, thus improving overall efficiency while ensuring image quality.
- Precision video editingWhile preserving the motion trajectory and spatial structure of the original live-action footage, high-precision style rewriting, scene redrawing, and character replacement are carried out to achieve controllable creation of "live-action skeleton + AI stylized overlay".
- Automatic duration predictionThe built-in Duration Head module reads the action description in the prompt and automatically determines the most suitable segment length, eliminating the need for creators to experiment repeatedly.
- Professional-grade output formatSupports native 4K HDR, RAW and EXR lossless output formats, and can be directly connected to professional film and television color grading, special effects compositing and post-production editing pipelines.
- Synchronous audio generation: Generate matching sound tracks for video synchronization using a standalone Audio VAE module.
- IC-LoRA Precision ControlIt provides three reference video control modes: Canny Edge, Depth, and Pose, ensuring that the generated results strictly follow the composition, spatial structure, or character movements of the reference material.
Technical Principles of LTX-2.5
- Gemma 4 12B Text EncoderLightricks' custom-designed large language model encoder for LTX fully preserves multi-subject relationships, action details, lighting descriptions, and camera orientation in complex and long cues, without losing clause information as cues become longer.
- Prompt EnhancerA lightweight auxiliary model that automatically expands short prompts input by the user into detailed, cinematic instructions, lowering the engineering threshold for prompt words and with extremely low computational overhead.
- Diffusion Video DecoderA new architecture that replaces the traditional VAE decoder is specifically optimized for fast-moving scenes, significantly reducing image smearing and artifacts, making faces clearer and text more readable.
- Audio VAEAn independent audio latent space encoder is responsible for mapping video content to matching audio representations, enabling the synchronous generation of video and sound.
- Diffusion Fidelity RenderingThe system employs a two-stage rendering strategy: "structure first, details later." The first stage constructs motion, composition, and shot framing within an 8× time-compressed latent space and adaptively generates high-fidelity keyframes. The second stage diffuses and renders the final video from both the structure and keyframes, resolving textures, materials, and facial details with pixel-level precision.
- Duration HeadThe prediction module, deployed before the diffusion process, automatically infers and outputs the most suitable video length by understanding the action semantics in the prompts, enabling the model to have a preliminary ability to perceive the narrative rhythm.
Follow us on WeChat and reply with "open source",join inAI open source project discussion group
How to use LTX-2.5
-
Select pathThe decision was made to use either the API cloud service or ComfyUI local deployment.
-
API AccessApply for an API Key on the LTX official website and select either Fast or Pro to call the API.
-
Installation NodeFor local deployment, ComfyUI needs to be updated to the latest version and installed via Manager.
ComfyUI-LTXVideoNode package. -
Download ModelSearch for "LTX-2.5" in the ComfyUI template browser, select the corresponding workflow, and download all model files with one click.
-
Configuration parametersSet the width and height to multiples of 32, the frame rate to multiples of 1+8, the frame rate to 24/25/48/50fps, and the base resolution to half of the target to enable 2× scaling.
-
Write promptsUse a coherent prose passage to describe the shot composition, scene lighting, character movements, camera movement, and audio, avoiding tag stacking or weighted syntax.
-
Execution generationClick to run. For text-based videos, upload the first frame directly. For image-based videos, upload both the first and last frames.
-
Precise controlAdvanced users can load IC-LoRA in Canny, Depth, or Pose modes and connect to reference videos to achieve composition or motion locking.
-
Model selectionUse a distilled model for rapid iteration, and a full dev model for the final output to obtain finer details.
The core advantages of LTX-2.5
- Narrative coherenceNative Multishot multi-camera generation capability, which can maintain a high degree of consistency in character appearance, environmental atmosphere, lighting conditions and sound in a single output, fundamentally solving the biggest pain point of shot continuity in AI plot creation.
- A leap in controllabilityPrecise video editing preserves the motion trajectory and spatial structure of the original live-action footage. Combined with the three IC-LoRA control modes of Canny/Depth/Pose, it enables a controllable production transformation from "blindly drawing cards" to "live-action skeleton + AI stylized overlay".
- Industrial-grade workflowNative RAW/HDR/EXR output can be directly connected to professional color grading, special effects compositing, and post-production editing pipelines without compression or transcoding, meeting the film and television industry's rigid demand for lossless materials.
- Adaptive efficiencyDiffusion Fidelity Rendering technology dynamically allocates rendering power according to scene complexity, enabling detailed rendering of complex dynamic shots and automatically saving resources for static and simple scenes, thus improving overall generation efficiency while ensuring pixel-level image quality.
LTX-2.5 project address
- Project official website:https://ltx.io/model/ltx-2-5
- HuggingFace model library:https://huggingface.co/Lightricks/LTX-2.5
Comparison of LTX-2.5 with similar competing products
| Comparison Dimensions | LTX-2.5 | Minimax H3 |
|---|---|---|
| Producer | Lightricks (Israel) | MiniMax (China) |
| Release time | August 2026 | August 2026 |
| Openness | The 22B weights are completely open source and can be deployed and fine-tuned locally. | Weighting is partially open in certain regions (excluding the US/UK/EU/Korea), and local fine-tuning is not supported. |
| Parameter size | 22B (Public) | Not disclosed |
| Maximum resolution | 4K (3840×2160) | 2K(fixed) |
| Maximum duration | 20 seconds | 10 seconds |
| Native Multi-Lens | support | Not supported |
| Synchronized audio | support | Not open |
| Input mode | Text / Image / Audio | Text / Image |
| Output format | 4K HDR, RAW, EXR lossless | Standard compression format |
| Controllability | Precise video editing + IC-LoRA (Canny/Depth/Pose) | Primarily relies on prompt words for driving |
Application scenarios of LTX-2.5
- Professional film and television post-productionNative 4K HDR, RAW, and EXR lossless output can be directly integrated into color grading, effects compositing, and editing pipelines, meeting the post-production needs of theatrical and streaming media, eliminating the nightmare of color grading compressed MP4.
- AI short dramas and comicsMultishot generates coherent narrative storyboards in a single output using multiple shots, and supports precise video editing with "real-life action skeleton + AI character replacement and stylized overlay," fundamentally solving the two major weaknesses of shot continuity and action controllability.
- Advertising and Brand ContentRapidly iterate on high-quality commercial videos, leveraging Prompt Enhancer to lower the engineering threshold for prompt words, with Distilled models supporting quick previews and complete model output for the final product, balancing efficiency and quality.
- Physics, AI and RoboticsPhysical AI Checkpoint provides embodied intelligence teams with a foundation of world models, helping robotic systems perceive the physical world and laws of motion, going beyond the realm of pure content creation.