> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs-dev.ltx.io/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs-dev.ltx.io/_mcp/server.

# LTX ComfyUI Nodes

> Reference LTX custom ComfyUI nodes used by LTX-2.3 and LTX-2.5 workflows, including Native HDR, SDR to HDR, Native Resolution tiled sampling, conditioning, audio, and VAE decoding.

This page documents key ComfyUI nodes that LTX ships in the [ComfyUI-LTXVideo](https://github.com/Lightricks/ComfyUI-LTXVideo) repository. Some are optional enhancements to existing workflows (faster prompt iteration, lower VRAM, finer guidance and IC-LoRA control) while others are required components of specific workflows, such as in-outpainting, text-to-audio, and tiled VAE decoding. Check each node's "When to use" notes for where it fits.

## Overview

**Gemma Text Encoding:**

* **GemmaAPITextEncode** - Free API-based text encoder that replaces the local Gemma and allows for reduced VRAM usage and faster runtimes
* **LTXVSaveConditioning** - Save text encodings to disk for reuse
* **LTXVLoadConditioning** - Load pre-saved text encodings

**Audio Identity (Dub-It):**

* **LTXVSetAudioRefTokens** - Attaches reference audio as conditioning tokens for speaker identity transfer

**Advanced Guidance:**

* **MultimodalGuider** - Independent control over audio and video guidance parameters
* **LTX Add Video IC-LoRA Guide Advanced** - Granular IC-LoRA strength control with global scaling and spatial masking

**In-Outpainting:**

* **LTXVInpaintPreprocess** - Prepares masked video input for inpainting/outpainting generation stages
* **LTXVLaplacianPyramidBlend** - Blends generated content with original video at mask boundaries

**Text-to-Audio:**

* **LTXVAudioOnlyModel** - Puts the joint AV transformer into audio-only mode for text-to-audio generation

**Quality Enhancement:**

* **LTXVNormalizingSampler** - Latent normalization to prevent overbaking and audio clipping

**Image Conditioning:**

* **LTXVImgToVideoConditionOnly** - Applies image conditioning to the first frames of a video latent

**Native HDR and SDR to HDR:**

* **LTXVSDRToHDRWorkingSpace** - Maps SDR or declared float RGB input into the ACEScct working space used by the LTX-2.5 SDR to HDR IC-LoRA
* **LTXVLoadEXRSequence** - Loads an EXR still or frame sequence and converts supported input into ACEScct for LTX-2.5 Native HDR workflows
* **LTXVVAEForceFloat32** - Forces the video VAE to float32 for LTX-2.5 Native HDR encode and decode
* **LTXVHDRDecodePostprocess** - Converts LTX-2.5 ACEScct or legacy LTX-2.3 LogC3 VAE output into a preview and scene-linear HDR output, with optional EXR writing
* **LTXVSaveHLG** - Encodes scene-linear HDR frames as a BT.2020/HLG 10-bit HEVC master

**Tiled Sampling:**

* **LTXVTiledFusionSampler** - Runs per-step overlapping IC-LoRA tiles on one shared latent canvas and noise field
* **LTXVGetTilingSizes** - Resolves tile, canvas, and output sizes for a Tiled Fusion workflow into model-legal geometry

**VAE Decoding:**

* **LTXVTiledVAEDecode** - Tiled VAE decode to reduce VRAM
* **LTXVSpatioTemporalTiledVAEDecode** - Space- and time-tiled decode for long or large video

---

## HDR and EXR nodes

Keep versioned HDR workflows and assets separate. The only cross-version exception documented by the current LTX-2.5 VFX graphs is the specific LTX-2.3 In/Outpainting IC-LoRA used by the HDR inpainting workflow.

### LTXVSDRToHDRWorkingSpace

Location: [`hdr_nodes.py`](https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/hdr_nodes.py)

**What it does**

Maps decoded SDR or declared float RGB frames into ACEScct for the LTX-2.5 SDR to HDR IC-LoRA. Apply the input transform before resizing so display-encoded sRGB values are linearized before interpolation.

**When to use**

Use this node in the LTX-2.5 SDR to HDR workflow before resize and IC-LoRA guide creation. Do not use it for Native HDR EXR ingest; use `LTXVLoadEXRSequence` for that path.

**Parameters**

| Parameter     | Type    | Default      | Description                                                                                                                                                            |
| ------------- | ------- | ------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `image`       | `IMAGE` | Required     | Source RGB frames.                                                                                                                                                     |
| `color_space` | Choice  | `srgb_gamma` | `srgb_gamma` for ordinary display-encoded sRGB video; `srgb` for scene-linear Rec.709/sRGB; `acescg` for scene-linear ACEScg; or `acescct` for an ACEScct passthrough. |

**Returns**

* `acescct` - ACEScct working-space frames for downstream resize and HDR IC-LoRA conditioning.

### LTXVLoadEXRSequence

Location: [`hdr_nodes.py`](https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/hdr_nodes.py)

**What it does**

Loads one EXR still or a naturally sorted folder of EXR frames, interprets the pixels in the declared source color space, and converts them to ACEScct for an LTX-2.5 native HDR workflow. It also reports sequence geometry and creates a silent audio object with matching duration for graph compatibility.

**When to use**

Use this node in the published native EXR image-to-video and inpainting workflows. Do not use it to treat an HLG, PQ, HDR10, MP4, or MOV file as native HDR input.

**Parameters**

| Parameter     | Type    | Default       | Description                                                                                                        |
| ------------- | ------- | ------------- | ------------------------------------------------------------------------------------------------------------------ |
| `path`        | String  | Required      | Path to one `.exr` file or a directory of `.exr` frames. Relative paths resolve under the ComfyUI input directory. |
| `color_space` | Choice  | `srgb_linear` | Declared source pixels: `srgb_linear`, `acescg`, or `acescct`.                                                     |
| `frame_rate`  | Float   | `24.0`        | Sequence rate from 1 to 120 fps. EXR folders do not contain container timing.                                      |
| `frame_start` | Integer | `0`           | Number of naturally sorted source frames to skip.                                                                  |
| `frame_cap`   | Integer | `0`           | Maximum frames to load. `0` loads all remaining frames.                                                            |
| `trim_to_8k1` | Boolean | `true`        | Trims the result to the longest valid `8k+1` prefix required by the video VAE.                                     |

**Returns**

* `images` - ACEScct frames.
* `frame_rate` - The configured sequence rate.
* `width`, `height`, and `frame_count` - Loaded sequence geometry.
* `audio` - Silent audio matching the loaded duration.

> **Warning**
>
> The loader reads RGB frames and requires every frame to have the same shape. Do not assume it preserves alpha, auxiliary channels, camera-RAW data, or arbitrary EXR metadata.

### LTXVVAEForceFloat32

Location: [`hdr_nodes.py`](https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/hdr_nodes.py)

**What it does**

Sets the supplied video VAE and its first-stage model to float32 for LTX-2.5 native HDR encode and decode.

**When to use**

Place this node immediately after the video VAE loader and route all native HDR encode, decode, and IC-LoRA guide operations through its output.

**Parameters**

| Parameter | Type  | Default  | Description                                |
| --------- | ----- | -------- | ------------------------------------------ |
| `vae`     | `VAE` | Required | Video VAE used by the native HDR workflow. |

**Returns**

* `VAE` - The same VAE instance after conversion to float32.

> **Warning**
>
> This node mutates the input VAE; it does not create an independent float32 copy. A parallel branch using the loader output shares the mutated model and can become execution-order dependent. Use one chain from the VAE loader through this node to every HDR consumer.

### LTXVHDRDecodePostprocess

Location: [`hdr_nodes.py`](https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/hdr_nodes.py)

**What it does**

Decompresses VAE-decoded HDR into scene-linear values, produces a Reinhard-tonemapped sRGB preview, and can write an EXR frame sequence. The `acescct` transfer supports the LTX-2.5 Native HDR and SDR to HDR paths. The `logc3` transfer is for the legacy LTX-2.3 HDR IC-LoRA.

**When to use**

Place this node after VAE Decode. Send `tonemapped` to an ordinary preview and `hdr_linear` to `LTXVSaveHLG` or another scene-linear HDR operation.

**Parameters**

| Parameter         | Type    | Default          | Description                                                                                       |
| ----------------- | ------- | ---------------- | ------------------------------------------------------------------------------------------------- |
| `image`           | `IMAGE` | Required         | VAE-decoded HDR frames in the selected working-space curve.                                       |
| `transfer`        | Choice  | `acescct`        | `acescct` for LTX-2.5 Native HDR and SDR to HDR; `logc3` only for the legacy LTX-2.3 HDR adapter. |
| `exposure`        | Float   | `0.0`            | Preview exposure from -10 to +10 EV. It does not change `hdr_linear` or saved EXR data.           |
| `save_exr`        | Boolean | `false`          | Writes an EXR sequence when enabled.                                                              |
| `exr_color_space` | Choice  | `acescct`        | `acescct`, scene-linear `acescg`, scene-linear `srgb_linear`, or legacy untagged `linear`.        |
| `output_dir`      | String  | `output/hdr_exr` | EXR directory. Relative paths resolve under the ComfyUI output directory.                         |
| `filename_prefix` | String  | `frame`          | Prefix for `<prefix>_XXXXX.exr` output.                                                           |
| `half_precision`  | Boolean | `true`           | Writes float16 EXR when enabled and float32 when disabled.                                        |

**Returns**

* `tonemapped` - SDR preview after exposure, Reinhard tonemapping, and sRGB encoding.
* `hdr_linear` - Scene-linear HDR. With `acescct`, this is ACEScg; with `logc3`, it uses the legacy Rec.709-style path.

### LTXVSaveHLG

Location: [`hdr_nodes.py`](https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/hdr_nodes.py)

**What it does**

Converts scene-linear HDR into Rec.2020 and writes a BT.2020/HLG 10-bit HEVC MP4. It can mux an optional audio input as AAC.

**When to use**

Connect the `hdr_linear` output from `LTXVHDRDecodePostprocess`, not the tonemapped preview. Match `linear_primaries` to the selected decode transfer.

**Parameters**

| Parameter          | Type    | Default        | Description                                                                            |
| ------------------ | ------- | -------------- | -------------------------------------------------------------------------------------- |
| `hdr_linear`       | `IMAGE` | Required       | Scene-linear HDR frames from `LTXVHDRDecodePostprocess`.                               |
| `frame_rate`       | Float   | `24.0`         | Output rate from 1 to 120 fps.                                                         |
| `filename_prefix`  | String  | `hdr/ltxv_hlg` | ComfyUI output prefix. The saved filename receives an `_hlg.mp4` suffix.               |
| `linear_primaries` | Choice  | `acescg`       | Use `acescg` for the LTX-2.5 ACEScct decode path or `rec709` for legacy LTX-2.3 LogC3. |
| `audio`            | `AUDIO` | Optional       | Audio track to encode and mux as AAC.                                                  |

**Returns**

This output node writes the HLG master and returns no workflow data. If the input width or height is odd, it removes the final column or row before 4:2:0 encoding.

---

## Gemma Text Encoding Nodes

### GemmaAPITextEncode

Location: [`gemma_api_conditioning.py`](https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/gemma_api_conditioning.py)

**What it does**

Encodes text prompts using LTX's free API endpoint, bypassing the need to load Gemma locally. This eliminates all local VRAM usage for text encoding and enables sub-second prompt encoding.

**Why use it**

Gemma's large memory footprint (requires loading/unloading from VRAM) can create a bottleneck on consumer hardware, particualry during prompt iteration. Every time you change a prompt, Gemma must be reloaded, adding significant time to the workflow. This node solves that problem by offloading text encoding to a free API endpoint.

**When to use**

Use this node when:

* Working on consumer GPUs with limited VRAM
* Multiple generations use different prompts

**Parameters**

* `api_key` - Your LTX API key
* `prompt` - The text prompt to encode
* `ckpt_name` - The LTX checkpoint file (used to extract model ID for encoding compatibility)

**Returns**

* `conditioning` - Encoded prompt conditioning ready for LTX generation

**Getting an API key**

1. Visit [console.ltx.io](https://console.ltx.io)
2. Sign up or log in
3. Generate a free API key
4. Copy the key into the node's `api_key` parameter

**Example workflow**

With API:

```
Encode via API → Generate → Change Prompt → Encode via API → Generate
```

Without API:

```
Load Gemma → Encode Prompt → Generate → Change Prompt → Reload Gemma → Encode → Generate
```

---

### LTXVSaveConditioning

Location: [`conditioning_saver.py`](https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/conditioning_saver.py)

**What it does**

Saves computed text conditioning to disk as a .safetensors file, allowing you to reuse the exact same conditioning across multiple workflow sessions without re-encoding.

**Why use it**

Useful when:

* You have a prompt that works well and want to preserve its exact encoding
* Running batch generations with identical conditioning
* Building reusable workflow templates with pre-encoded prompts
* Working offline without API access

**When to use**

Use this node when:

* You want to lock in a specific prompt's encoding
* Multiple workflow sessions will use the same conditioning
* You need reproducible conditioning across different machines
* Building libraries of validated prompts

**Parameters**

* `conditioning` - The conditioning to save (from any text encoder or the API node)
* `filename` - Base filename (without extension)
* `dtype` - Precision for storage: "bfloat16" or "float16"

**Returns**

* UI notification showing saved filename and file size

**Output location**

Files are saved to: `ComfyUI/models/embeddings/`

**Storage**

Files are stored as .safetensors using the selected numerical precision:

* bfloat16: Higher precision, more commonly used
* float16: Alternative representation, minimal practical difference

---

### LTXVLoadConditioning

Location: [`conditioning_loader.py`](https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/conditioning_loader.py)

**What it does**

Loads precomputed conditioning from `ComfyUI/models/embeddings/`, bypassing text encoding. In addition to conditioning saved by `LTXVSaveConditioning`, the node accepts pipeline scene embeddings containing `video_context` or `video_prompt_embeds` and optional audio context.

**Why use it**

Use saved conditioning to avoid repeated text encoding and preserve an exact conditioning tensor. The pipeline-embedding path also lets the LTX-2.5 SDR to HDR workflow use its fixed video and audio scene conditioning without an ordinary free-text prompt.

**When to use**

Use this node when:

* Reusing conditioning saved in previous sessions
* Running batch workflows with preset prompts
* Working offline without API access
* You need bit-perfect conditioning reproducibility
* Running the LTX-2.5 SDR to HDR workflow with its matching fixed scene-embedding file

**Parameters**

* `file_name` - The `.safetensors` file to load from the embeddings folder
* `device` - Where to load the conditioning: "cpu" or "gpu"

**Returns**

* `conditioning` - Loaded conditioning ready for generation

If the selected pipeline file contains both video and audio contexts, the node checks their leading dimensions and concatenates them into the conditioning layout expected by the audio-video model. A file with no recognized conditioning keys fails instead of returning empty conditioning.

**Device selection**

* **cpu**: Loads to system RAM (slower but works on any system)
* **gpu**: Loads directly to VRAM if available (faster for generation)

**Workflow integration**

Pair with LTXVSaveConditioning to create prompt libraries:

1. Create and refine prompts with text encoder or API
2. Save successful conditioning with LTXVSaveConditioning
3. Load instantly in future sessions with LTXVLoadConditioning

For the LTX-2.5 SDR to HDR path, use `ltx-2.5-22b-ic-lora-sdr-to-hdr-scene-emb.safetensors`. Do not substitute the legacy LTX-2.3 HDR scene embeddings.

---

## Advanced Model Guidance

### MultimodalGuider

Location: [`multimodal_guider.py`](https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/guiders/multimodal_guider.py)

**What it does**

Provides independent, per-modality control over guidance parameters for audio and video. This is an extension of Classifier-Free Guidance (CFG) that allows you to separately control prompt adherence, artifact reduction, and cross-modal synchronization for each modality.

**Why use it**

Standard guidance treats audio and video as a single unit. When you increase guidance to improve video quality, it affects audio synchronization. When you fix synchronization, your visual style can break. The MultimodalGuider decouples these controls, letting you tune video guidance independently from audio guidance without trade-offs.

**When to use**

Use this node when:

* You need different guidance strengths for audio vs video
* Video quality needs to be prioritized over tight audio sync (or vice versa)
* You want to prevent the common issue where fixing synchronization breaks visual style
* You need fine-grained control over cross-modal attention

**How it works**

The guider can make up to four separate model inference calls per step:

1. **Positive conditioning** - Your prompt
2. **Negative conditioning** - Your negative prompt (for CFG)
3. **Perturbed conditioning** - Degraded version (for STG artifact reduction)
4. **Modality-isolated conditioning** - Each modality without cross-attention (for sync control)

By combining these strategically, you get independent control over:

* CFG strength per modality (prompt adherence)
* STG strength per modality (artifact reduction)
* Cross-modal attention strength (synchronization tightness)
* Step skipping per modality (performance optimization)

**Parameters**

* `model` - The LTX model to apply guidance to
* `positive` - Positive conditioning
* `negative` - Negative conditioning
* `parameters` - A GUIDER\_PARAMETERS object containing per-modality settings
* `skip_blocks` - Comma-separated list of transformer blocks to skip for STG

**GUIDER\_PARAMETERS structure**

The parameters object exposes three independent guidance controls, each configurable per modality (audio and video separately):

**1. CFG Guidance (cfg > 1)**

Controls prompt adherence and semantic accuracy. Pushes the model toward the positive prompt and away from the negative prompt.

* **When to increase:** When visual style or object fidelity matters most
* **Effect:** Stronger prompt following, more accurate semantic content
* **Configurable per modality:** Yes

**2. Spatio-Temporal Guidance (stg > 0)**

Reduces artifacts by pushing the model away from a degraded, perturbed version of itself. Prevents breakup of rigid objects. Based on the [STG technique](https://arxiv.org/abs/2411.18664).

* **When to increase:** If you see structural artifacts or object breakup
* **Effect:** Fewer visual artifacts, more stable structures
* **Configurable per modality:** Yes

**3. Cross-Modal Guidance (modality\_scale > 1)**

Controls synchronization between audio and video. Pushes the model away from versions where modalities ignore each other.

* **When to adjust:** To balance synchronization versus natural motion
* **Higher values:** Tighter alignment (perfect for lip-sync or rhythmic action)
* **Lower values:** Looser, more natural coupling
* **Configurable per modality:** Yes

**Additional Per-Modality Parameters**

* **`skip_step`** - Periodically skip diffusion steps for this modality
  * `0`: No skipping
  * `1`: Skip every other step
  * `2`: Skip two out of every three steps
  * Use for performance optimization

* **`rescale`** - Normalization after applying CFG, STG, and cross-modal guidance
  * `0`: No normalization
  * `1`: Full renormalization to match the norm of the positive-prompt prediction
  * `0-1`: Partial normalization
  * Especially helpful for preventing oversaturation when using high CFG or STG values

* **`perturb_attn`** - Boolean controlling whether the perturbed model is perturbed for this modality during STG. Normally set to `True`.

* **`cross_attn`** - Boolean controlling whether cross-attention layers from this modality to the other modality are active. Normally set to `True`.

**Returns**

* `guider` - Configured guider ready for sampling

**Use cases**

**Use Case 1: Prioritize video quality, loose audio sync**

* Video: High CFG, moderate STG, low modality scale
* Audio: Low CFG, low STG, low modality scale
* Result: Beautiful video, audio follows general mood but not frame-locked

**Use Case 2: Tight lip-sync for dialogue**

* Video: Moderate CFG, moderate STG, high modality scale
* Audio: Moderate CFG, low STG, high modality scale
* Result: Audio and video tightly synchronized, good for speaking

**Use Case 3: Performance optimization**

* Video: Process every step
* Audio: Skip every other step (skip\_step = 1)
* Result: 2x faster generation with minimal audio quality impact

**Integration with other nodes**

* Works with all LTX sampler nodes
* Can be combined with latent normalization for additional quality control
* Essential for looping sampler workflows

---

### LTX Add Video IC-LoRA Guide Advanced

Location: [`iclora.py`](https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/iclora.py)

**What it does**

Applies an IC-LoRA control adapter with granular strength control, replacing the fixed 1.0 strength behavior of the standard IC-LoRA node. Allows global strength adjustment and optional spatial/spatiotemporal masking. This is also the guide node used in the [In-Outpainting workflow](/open-source-model/feature-guides/editing-effects/in-outpainting), where it passes mask information through the IC-LoRA pipeline so the model knows which regions to generate and which to preserve.

**Why use it**

The standard IC-LoRA workflow applies control at full strength everywhere, which can over-constrain generation. This node lets you dial in exactly how much influence the control signal has, and where. It is also required for mask-aware IC-LoRA workflows like In-Outpainting.

**When to use**

Use this node when:

* You want softer, less rigid IC-LoRA control
* You need IC-LoRA to apply only to specific regions of the frame
* You want to blend IC-LoRA control with free generation
* You're combining multiple control types and need to balance their influence
* You're running an [In-Outpainting](/open-source-model/feature-guides/editing-effects/in-outpainting) workflow that requires mask-aware conditioning

**Parameters**

* `attention_strength` (float, 0.0-1.0) — Global scaling factor for IC-LoRA cross-attention scores. Default: 1.0
* `attention_mask` (MASK, optional) — Spatial (H×W) or spatiotemporal (T×H×W) mask multiplied with attention\_strength

The node exposes additional widget parameters used by mask-aware workflows like [In-Outpainting](/open-source-model/feature-guides/editing-effects/in-outpainting). For those workflows, use the pre-configured defaults from the workflow JSON.

**Returns**

* Model with IC-LoRA applied at the specified strength/mask configuration

---

## Quality Enhancement

### LTXVNormalizingSampler

Location: [`easy_samplers.py`](https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/easy_samplers.py)

**What it does**

A specialized sampler that applies statistical normalization to latents during generation to prevent overbaking (oversaturation) and audio clipping issues.

**Why use it**

Without normalization, latent values can drift into problematic ranges during the denoising process. This causes:

* Oversaturated, "overbaked" visual outputs with crushed colors
* Audio clipping and distortion
* Inconsistent quality across different prompts or settings

The NormalizingSampler keeps latent statistics in optimal ranges throughout generation, dramatically improving output quality.

**When to use**

Use this node when:

* You see oversaturated, "overbaked" visual outputs
* Audio has clipping or distortion artifacts
* Output quality varies unpredictably between generations
* Using high guidance values that tend to cause oversaturation

**How it works**

The sampler monitors latent statistics during the denoising process and applies normalization to keep values within target ranges. This is done using percentile-based statistics (excluding extreme outliers) to prevent both overbaking and excessive normalization.

**Key benefits**

* Prevents oversaturated, "overbaked" visual outputs
* Eliminates audio clipping artifacts
* More consistent quality across generations
* Works automatically - no manual tuning required
* Especially effective with high guidance values

**Integration**

This is a drop-in replacement for standard samplers in LTX workflows. It maintains full compatibility with:

* All guider nodes (including MultimodalGuider)
* Text and image conditioning
* LoRA and IC-LoRA workflows

**Performance impact**

Minimal - the normalization adds negligible computational overhead while significantly improving output quality.

---

## Audio Identity

This node provides speaker identity conditioning for the [Dub-It IC-LoRA](/open-source-model/feature-guides/audio/dub-it-beta) two-stage dubbing pipeline.

### LTXVSetAudioRefTokens

Location: [`iclora.py`](https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/iclora.py)

**What it does**

Attaches an audio latent as `ref_audio` tokens on conditioning for speaker identity transfer. Also outputs a `frozen_audio` copy with `noise_mask=0`, ensuring Stage 1 audio passes through Stage 2 unchanged without needing a mask-by-time node.

**Why use it**

The Dub-It pipeline needs to preserve the original speaker's voice across both generation stages. This node handles two things in one step: it gives the model the speaker's audio identity as reference tokens, and it freezes the audio latent so it carries forward unchanged into Stage 2.

**When to use**

Used once per stage in the Dub-It two-stage pipeline. Stage 1 receives the VAE-encoded reference audio; Stage 2 receives the Stage 1 audio output.

**Parameters**

* `conditioning` - The text conditioning to attach reference tokens to
* `audio_latent` - The audio latent to use as reference (from VAE encode in Stage 1, or from Stage 1 output in Stage 2)

**Returns**

* `conditioning` - Conditioning with `ref_audio` tokens prepended
* `frozen_audio` - Audio latent with zero noise mask for pass-through

---

## In-Outpainting

These nodes are used in the [In-Outpainting workflow](/open-source-model/feature-guides/editing-effects/in-outpainting) for extending or filling regions of existing video. They work alongside `LTX Add Video IC-LoRA Guide Advanced` (documented above) to provide mask-aware video generation.

### LTXVInpaintPreprocess

Location: [`vanish_nodes.py`](https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/vanish_nodes.py)

**What it does**

Prepares masked video input for inpainting or outpainting generation. Takes the source video and its mask and formats them for the sampler. Used twice in the two-stage pipeline — once at base resolution (stage 1) and once at upscaled resolution (stage 2).

**When to use**

Required at each generation stage in the In-Outpainting workflow to prepare the masked input before sampling.

**Parameters**

This node has no configurable widget parameters. It accepts two inputs:

* `images` (IMAGE) — The source video frames
* `mask` (MASK) — The generation mask

---

### LTXVLaplacianPyramidBlend

Location: [`pyramid_blending.py`](https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/pyramid_blending.py)

**What it does**

Blends generated video with original content using Laplacian pyramid blending, producing seamless transitions at mask boundaries. Applied after each generation stage to merge new content with the preserved original video, preventing visible seams between generated and original regions.

**When to use**

Required after each generation stage in the In-Outpainting workflow. This is the primary node for controlling output quality at mask boundaries — if you see seams or color mismatches, this is the node to adjust.

**Parameters**

* `dilation` (int) — Controls how far the blending extends beyond the mask edge. The most impactful parameter for boundary quality. Workflow defaults: **5** (stage 1) and **6** (stage 2). Higher values extend the blend region further, which helps when the boundary area contains low-frequency content (smooth gradients, sky, etc.).

---

## Text-to-Audio

### LTXVAudioOnlyModel

Location: [`audio_only.py`](https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/audio_only.py)

**What it does**

Configures the LTX joint audio-video transformer to run in audio-only mode, disabling all video computation and cross-modal attention. This enables text-to-audio generation without producing video output — the model generates audio directly from a text prompt.

**Why use it**

LTX is a single transformer that processes audio and video together. By default, both modalities run on every forward pass, including cross-attention between them. When you only need audio, this node eliminates all video overhead — the video stream is never computed, and the audio is denoised independently.

**When to use**

Use this node when you want to generate audio from text without any video. It is the core node in the [Text-to-Audio workflow](/open-source-model/usage-guides/text-to-audio).

**How it works**

The node sets three `transformer_options` flags on the model:

* `run_vx = False` — Skips the video stream's self-attention and feedforward layers entirely
* `a2v_cross_attn = False` — Skips audio→video cross-attention
* `v2a_cross_attn = False` — Skips video→audio cross-attention

This is equivalent to running the reference pipeline with `video=None`. The audio pathway runs normally while all video compute is eliminated.

> **Note**
>
> The model architecture still expects a video input in the latent. You must provide a minimal dummy video latent (64×64, 1 frame via `EmptyLTXVLatentVideo`) concatenated with the audio latent via `LTXVConcatAVLatent`. With video computation disabled, the dummy latent is never attended to and adds negligible cost.

**Parameters**

* `model` (MODEL) — The LTX model to configure for audio-only mode
* `skip_video_compute` (boolean, default: True) — Toggles the `run_vx` flag. When enabled, the video stream's self-attention and FFN are skipped entirely. The audio output is identical regardless of this setting, because audio↔video cross-attention is always disabled in this mode — this flag only controls whether the video compute is also skipped for performance.

**Returns**

* `model` (MODEL) — The same model configured for audio-only generation

---

## Image Conditioning

### LTXVImgToVideoConditionOnly

Location: [`latents.py`](https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/latents.py)

**What it does**

Applies image conditioning to the first frames of an existing video latent. It encodes the supplied image with the VAE, resizes it to match the latent's dimensions, and writes it into the latent along with a noise mask that controls how strongly the image is enforced. No sampling happens here — the node only prepares the conditioned latent for a downstream sampler.

**When to use**

Use it in image-to-video and IC-LoRA video-to-video workflows to pin the opening frame(s) to a supplied image before sampling. It appears in many of the shipped example workflows. The `bypass` toggle lets you disable the conditioning without rewiring the graph — handy in workflows where first-frame conditioning is optional.

**Parameters**

* `vae` (VAE) — VAE used to encode the image into latent space
* `image` (IMAGE) — The image to condition on (auto-resized to the latent's dimensions)
* `latent` (LATENT) — The video latent to apply conditioning to
* `strength` (float, default: 1.0, range 0.0–1.0) — How strongly the image is enforced on the first frames, via the noise mask. 1.0 is full conditioning
* `bypass` (boolean, optional, default: False) — When True, passes the latent through unchanged (conditioning disabled)

**Returns**

* `latent` (LATENT) — The latent with image conditioning applied

---

## Native Resolution tiled sampling

Native Resolution divides diffusion-model work across a larger latent canvas with Tiled Fusion. It is separate from the VAE decoding nodes in the next section, which operate only after sampling is complete.

### LTXVTiledFusionSampler

**What it does**

Runs each denoising step on overlapping spatial crops, then Gaussian-blends the stepped crops back into one shared latent canvas. All tiles participate in one denoising trajectory and use one noise field.

This is distinct from [`LTXVTiledSampler`](https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/tiled_sampler.py), which samples tiles independently, and from tiled VAE decoding, which runs only after sampling.

**When to use**

Use the node for an IC-LoRA video-to-video task whose output canvas is larger than the adapter's trained spatial window. This can include a full-HD output when the adapter was trained on a smaller tile. Match the tile dimensions to the trained window documented for that adapter.

Full HD, 4K, and 8K are available in the example workflows, but the size selector is not a performance guarantee. Use only a model and IC-LoRA explicitly matched to the selected workflow, and test the target resolution on your hardware.

**Required graph contract**

* Connect `positive`, `negative`, and `latents` from the IC-LoRA guide node. A bare empty video latent does not contain the guide frames and noise mask required by the sampler.
* Keep `use_tiled_encode` set to `false` on the guide node. Spatially tiled guide encoding can imprint a grid into the conditioning.
* Do not add `LTXVCropGuides`; Tiled Fusion crops appended guide data for each tile.
* Supply descending `sigmas` that end at zero.
* Supply a `SAMPLER` from `KSamplerSelect` or another compatible node that exposes `sampler_function`.

**Parameters**

| Parameter              | Default  | Description                                                                                                                                                                                                                                    |
| ---------------------- | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `model`                | Required | Diffusion model used for each tile.                                                                                                                                                                                                            |
| `positive`, `negative` | Required | Conditioning returned by the IC-LoRA guide node.                                                                                                                                                                                               |
| `latents`              | Required | Latent returned by the IC-LoRA guide node.                                                                                                                                                                                                     |
| `sigmas`               | Required | Descending denoising schedule ending at zero.                                                                                                                                                                                                  |
| `sampler`              | Required | Compatible discrete-step sampler object.                                                                                                                                                                                                       |
| `seed`                 | `42`     | Seed for the shared full-canvas noise field.                                                                                                                                                                                                   |
| `cfg`                  | `1.0`    | Classifier-free guidance scale. Use the value specified by the selected workflow.                                                                                                                                                              |
| `tile_width`           | `1024`   | Tile width in pixels, in multiples of 32. Match the adapter's trained envelope.                                                                                                                                                                |
| `tile_height`          | `576`    | Tile height in pixels, in multiples of 32. Match the adapter's trained envelope.                                                                                                                                                               |
| `overlap_frac`         | `0.5`    | Fraction of spatial overlap. Values below `0.5` can produce a periodic grid on structured content.                                                                                                                                             |
| `blend_var`            | `0.05`   | Variance of the Gaussian blend weights.                                                                                                                                                                                                        |
| `grid_cycle`           | `1`      | Cycles shifted spatial grids across steps. `1` is a fixed grid; the maximum is `4`.                                                                                                                                                            |
| `tile_frames`          | `0`      | Temporal window in pixel frames (97 for LTX). `0` processes one extent. For longer clips, set the same value on both this node and the guide node with `use_streaming` on, so each window is re-encoded fresh; the example workflows use `97`. |
| `vae`                  | Optional | Supplies temporal and spatial downscale factors; otherwise they are read from the diffusion model.                                                                                                                                             |
| `canvas_device`        | `auto`   | Stores full-canvas tensors on `auto`, `gpu`, or `cpu`.                                                                                                                                                                                         |

**Sampler restrictions**

* Discrete step methods such as `euler`, `heun`, and `dpm_2` are the intended starting point.
* History-based methods such as `dpmpp_2m` and `lms` lose their multi-step memory because fusion owns the outer denoising loop.
* Adaptive samplers cannot fuse between steps.
* Ancestral methods add independent noise during individual tile steps and can create seams.

**Additional constraints**

* Batch size is limited to one.
* Downscaled guide latents whose spatial grid does not match the target canvas are rejected; reattach the guide at factor one for this sampler.
* Temporal windowing requires a guide encoded freshly for each window (`use_streaming` on the guide node with a matching `tile_frames`).
* Overlap and Gaussian blending reduce boundary risk but do not guarantee seamless output.

**Returns**

* `output` (`LATENT`) - The fused full-canvas latent for downstream VAE decoding.

### LTXVGetTilingSizes

**What it does**

Resolves tile and output sizes—and, for two-stage workflows, the stage-1 canvas—from named presets. It emits model-legal geometry: spatial dimensions that are multiples of 32 and a frame count of the form `8k+1`. It sits in the **Preprocess** subgraph of the Tiled Fusion examples and drives the canvas resize. The single-stage Upscale graph exposes only `tile_size` and `output_size`; the two-stage Native 4K/8K graph also exposes `initial_canvas_size`.

The named presets are workflow conveniences, not adapter compatibility settings. If an adapter's model card specifies a trained tile or canvas size that is not represented by a preset, set the custom dimensions inside **Preprocess** rather than substituting the nearest preset.

**Parameters**

| Parameter             | Example        | Description                                                                     |
| --------------------- | -------------- | ------------------------------------------------------------------------------- |
| `width`, `height`     | `1920`, `1080` | Source dimensions used to derive the resolved geometry.                         |
| `frame_count`         | `97`           | Requested frame count; snapped to the nearest valid `8k+1`.                     |
| `tile_size`           | `HD`           | Spatial tile preset: `qHD`, `HD`, or `FullHD`.                                  |
| `initial_canvas_size` | `FullHD`       | Stage-1 composition canvas preset, used by the two-stage Native 4K/8K workflow. |
| `output_size`         | `FullHD`       | Output preset: `FullHD`, `4K`, or `8K`.                                         |

**Returns**

* Resolved tile, canvas, and output geometry for the fusion sampler and the canvas resize.

---

## VAE Decoding

These nodes decode a video latent to pixels in tiles to reduce peak VRAM during the decode step, at the cost of slightly slower decoding. `LTXVTiledVAEDecode` is also explained in context in the [Two-Stage Distilled guide](/open-source-model/usage-guides/two-stage-generation).

### LTXVTiledVAEDecode

Location: [`tiled_vae_decode.py`](https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/tiled_vae_decode.py)

**What it does**

Decodes a video latent in overlapping spatial tiles rather than all at once, lowering peak VRAM usage during decode. The frame is split into a grid of tiles that are decoded separately and blended back together.

**When to use**

The most common decode node in the example workflows. Reach for it whenever a full-frame VAE decode runs out of memory; the defaults work for most hardware, and if you still hit OOM during decode, increase the tile count.

**Parameters**

* `vae` (VAE) — The VAE to decode with
* `latents` (LATENT) — The video latent to decode
* `horizontal_tiles` (int, default: 1, range 1–6) — Number of tiles across the width
* `vertical_tiles` (int, default: 1, range 1–6) — Number of tiles down the height
* `overlap` (int, default: 1, range 1–8) — Overlap between tiles (in latent units) to avoid visible seams
* `last_frame_fix` (boolean, default: False) — Repeats and then discards the last frame to work around a last-frame decode artifact
* `working_device` (`cpu` / `auto`, optional, default: `auto`) — Device used for decoding; `auto` matches the latents
* `working_dtype` (`float16` / `float32` / `auto`, optional, default: `auto`) — Precision used for decoding; `auto` matches the latents

**Returns**

* `image` (IMAGE) — The decoded video frames

---

### LTXVSpatioTemporalTiledVAEDecode

Location: [`tiled_vae_decode.py`](https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/tiled_vae_decode.py)

**What it does**

Extends tiled decoding across time as well as space — tiling the latent both spatially and along the temporal (frame) dimension. For long or large videos where spatial tiling alone still won't fit the decode in memory.

**When to use**

Use it for long or high-resolution clips where `LTXVTiledVAEDecode` still runs out of VRAM. It appears in several example workflows, often nested inside the low-VRAM subgraphs.

**Parameters**

* `vae` (VAE) — The VAE to decode with
* `latents` (LATENT) — The video latent to decode
* `spatial_tiles` (int, default: 4, range 1–8) — Number of spatial tiles, applied both horizontally and vertically
* `spatial_overlap` (int, default: 1, range 0–8) — Overlap between spatial tiles, in latent frames
* `temporal_tile_length` (int, default: 16, range 2–1000) — Length of each temporal tile in latent frames, including the overlap region
* `temporal_overlap` (int, default: 1, range 0–8) — Overlap between temporal tiles, in latent frames
* `last_frame_fix` (boolean, default: False) — Repeats and then discards the last frame after decoding
* `working_device` (`cpu` / `auto`, default: `auto`) — Device used for decoding; `auto` matches the latents
* `working_dtype` (`float16` / `float32` / `auto`, default: `auto`) — Precision used for decoding; `auto` matches the latents

**Returns**

* `image` (IMAGE) — The decoded video frames

---

## Installation & Usage

### Installation

All nodes are available in the [ComfyUI-LTXVideo](https://github.com/Lightricks/ComfyUI-LTXVideo) repository.

**Via ComfyUI Manager (Recommended):**

1. Open ComfyUI Manager
2. Search for "ComfyUI-LTXVideo"
3. Click **Update** (if already installed) or **Install**
4. Restart ComfyUI

**Manual Update:**

```bash
cd ComfyUI/custom_nodes/ComfyUI-LTXVideo
git pull origin master
pip install -r requirements.txt
```

After installation/update, restart ComfyUI. The new nodes will appear under the **"Lightricks"** category.

## Troubleshooting

### GemmaAPITextEncode

**"Invalid API key" error**

* Verify your API key is correct
* Regenerate a new key at console.ltx.io
* Ensure no extra spaces in the API key field

**"Cannot identify the text encoder" error**

* Your checkpoint file may be missing metadata
* Ensure you're using an official LTX model
* Ensure the ckpt\_name field in the node matches the filename of the model loaded in your Checkpoint Loader

**Timeout errors**

* Check your internet connection
* The API may be experiencing high load

### LTXVSaveConditioning / LTXVLoadConditioning

**File not found**

* Ensure .safetensors extension is not doubled
* Check that files are in `ComfyUI/models/embeddings/`
* Ensure you are not mixing up the `models/embeddings/` folder with `models/text_encoders/`

**Out of memory when loading**

* CPU / GPU memory management is key to avoiding OOM errors

### MultimodalGuider

**No quality improvement vs standard guider**

* Ensure you have connected two GuiderParameters nodes (one for audio, one for video)
* High values can break the generation. Suggested baseline for balanced speed and consistency is Modality: 1 and Skip Step: 1

**Generation is slower than expected**

* NOTE: This node uses CFG > 1, which inherently makes generation slower
* Use Skip Step: 1 for increased speed, reduce this value if artifacts appear

### LTXVNormalizingSampler

**Still seeing overbaking**

* This is a sampler, not a post-process — ensure you have swapped out the SamplerCustomAdvanced node
* Try combining with lower guidance values
* Consider if your prompt or conditioning is the root cause

**Quality seems worse**

* Use this node ONLY for the first sampling stage. Revert to a standard sampler for the second/upscale stage
* Do not use this sampler for inpainting, video extension, or any workflow using masks. It may break the context audio
* Ensure you are using the Distilled model with the standard 8-step manual sigma schedule. This node is NOT tuned for the full model
* Normalization helps most with problematic outputs (clipping/saturation). If your generation is already clean, this node may introduce unnecessary noise. Test side-by-side with the same seed