> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs-dev.ltx.io/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs-dev.ltx.io/_mcp/server.

# Two-Stage Generation

> Run the advanced LTX-2.5 two-stage ComfyUI workflow: base generation followed by a 2× upscale-and-refine pass.

This guide walks you through the **Two-Stage Distilled** workflow: a sample ComfyUI workflow from LTX that generates video with synchronized audio using a two-pass process: first at low resolution, then upscaled and refined at full resolution. It supports both Text-to-Video and Image-to-Video in a single workflow.

Compared to the default ComfyUI templates covered in the [Text-to-Video](/open-source-model/usage-guides/text-to-video) and [Image-to-Video](/open-source-model/usage-guides/image-to-video) beginner guides, this workflow exposes more of the pipeline and makes a good starting point for your own custom workflows.

If you haven't used the default templates yet, start there first. This guide assumes you're comfortable prompting, generating, and iterating in ComfyUI.

## What's Different from the Default Templates

The default templates get you generating quickly with minimal setup. This workflow uses the same distilled LTX-2.5 checkpoint and two-stage architecture, but exposes settings the templates handle automatically: the per-stage sigma schedules, tiled VAE decode, and the option to offload text encoding to the LTX API.

## Step-by-Step Guide

> **Note**
>
> Check the [system requirements](/open-source-model/getting-started/system-requirements) to make sure your hardware can run this workflow.

### 1. Download and Load the Workflow

Download the [Two-Stage Distilled workflow JSON](https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/example_workflows/2.5/LTX-2.5_T2V_I2V_Two_Stage_Distilled.json) and drag it into ComfyUI.

### 2. Install Custom Nodes and Download Models

This workflow requires **ComfyUI-LTXVideo**. Open the **Workflow Overview** panel after loading it; if any other custom nodes are missing, it lists and installs them.

Download the LTX-2.5 weights from the [LTX-2.5 HuggingFace repository](https://huggingface.co/Lightricks/LTX-2.5) (click **Agree and Access** on first download):

| File                                                      | Description                                                                   | Placement                               |
| --------------------------------------------------------- | ----------------------------------------------------------------------------- | --------------------------------------- |
| `ltx-2.5-22b-distilled-transformer-bf16.safetensors`      | Distilled LTX-2.5 transformer (loaded via **UNETLoader**)                     | `ComfyUI/models/diffusion_models/`      |
| `gemma4-12b-with-proj-ltx-2.5-bf16.safetensors`           | Text encoder (Gemma 4 12B)                                                    | `ComfyUI/models/text_encoders/`         |
| `gemma4_e2b_it_bf16.safetensors`                          | Prompt enhancer (Gemma 4 E2B); from `Comfy-Org/gemma-4`, not the LTX-2.5 repo | `ComfyUI/models/text_encoders/`         |
| `ltx-2.5-video-vae-bf16.safetensors`                      | Video VAE                                                                     | `ComfyUI/models/vae/`                   |
| `ltx-2.5-audio-vae-bf16.safetensors`                      | Audio VAE                                                                     | `ComfyUI/models/vae/`                   |
| `ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors` | Spatial upscaler (2×)                                                         | `ComfyUI/models/latent_upscale_models/` |

### 3. Choose Text-to-Video or Image-to-Video

In the **Input Parameters** subgraph, the **use image input** toggle switches modes: off = Text-to-Video (generate from the prompt), on = Image-to-Video (your image conditions the first frame). For Image-to-Video, load your image in the **LoadImage** node; **img\_strength** and **img\_compression** control how strongly it conditions the result.

### 4. Write Your Prompt

Write your prompt in the positive prompt field. For Text-to-Video, describe the full scene — setting, characters, camera movement, audio. For Image-to-Video, focus on motion, action, and audio, since the image supplies the visual context. Enable **enhance positive prompt** to expand a short prompt with the Gemma 4 enhancer (the negative prompt is never enhanced). See the [Prompting Guide](/open-source-model/usage-guides/prompting-guide) for additional tips.

### 5. Set Frame Rate and Duration

Set the **fps** and **duration in seconds**. The final frame count must be `1 + a multiple of 8`, so the actual duration is computed from fps × requested seconds and may differ slightly. Set **video width** and **video height** in the Preprocess subgraph — these are the Stage-1 dimensions; the pipeline upscales 2× in Stage 2, so the final output is twice these dimensions.

> **Warning**
>
> Higher resolution and longer duration need more VRAM. If you hit memory issues, reduce the resolution or duration, increase the VAE decode tile count, or enable API text encoding (below).

### 6. Generate

**Stage 1** generates the video and synchronized audio at the base resolution on the distilled checkpoint, using a short fixed sigma schedule. **Stage 2** upscales the Stage-1 result 2× and runs a few more steps to refine detail (re-using your image for conditioning if you supplied one). Click **Run** to generate.

### 7. Review and Iterate

The output is saved as an MP4 with synchronized audio. To iterate: adjust the prompt, change the duration, switch T2V/I2V, or try different resolutions.

## Customization Options

### Tiled VAE Decode

The tiled VAE decode step splits decoding into tiles to reduce peak VRAM at the cost of slightly slower decoding. Adjust tile count and overlap only if you hit memory issues during decode.

### CFG

Both stages use CFG 1. The distilled model bakes guidance into distillation, so raising CFG doesn't improve output the way it would with a standard diffusion model, and adds overhead. If you experiment, stay in the 1.0–1.5 range.

### API Text Encoding

To save VRAM, offload text encoding and enhancement to the LTX API by editing the Input Parameters subgraph. Remove the bypass, and use the "(via api)" outputs. You'll need an API key from the [LTX API Console](https://console.ltx.io).

## Using LoRAs

The official workflow does not include a standard-LoRA node. For a compatible standard LoRA, manually insert ComfyUI's `LoraLoaderModelOnly` after the shared **Model** output, then route its modified `MODEL` output to both the Stage 1 and Stage 2 samplers. Do not apply it to only one stage. See the [LoRA guide](/open-source-model/usage-guides/lo-ra).

## Python

Two-stage generation is also available through the PyTorch API for programmatic use and custom pipelines. See the [PyTorch API documentation](/open-source-model/integration-tools/pytorch-api).