> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs-dev.ltx.io/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs-dev.ltx.io/_mcp/server.

# Inpainting and Outpainting

> Fill masked regions or extend the frame of an existing video with the LTX Inpainting and Outpainting IC-LoRA workflows.

This guide covers the Inpainting and Outpainting workflows: sample ComfyUI workflows from LTX that fill or extend regions of an existing video using IC-LoRA conditioning. **Outpainting** generates new content beyond the original frame boundaries (making a video wider or taller); **inpainting** fills masked regions within the frame (removing or replacing objects). Both use the same In-Outpainting IC-LoRA with different mask configurations, and both run a two-stage pipeline that blends generated content seamlessly with the original.

## Prerequisites

This guide assumes you're familiar with ComfyUI basics and IC-LoRA workflows. If you're new to IC-LoRAs, start with the [IC-LoRA Guide](/open-source-model/usage-guides/ic-lo-ra).

## Model Files

Download the LTX-2.5 weights from the [LTX-2.5 HuggingFace repository](https://huggingface.co/Lightricks/LTX-2.5) (click **Agree and Access** on first download). The In-Outpainting IC-LoRA is reused from LTX-2.3.

| File                                                      | Description                                                                   | Placement                               |
| --------------------------------------------------------- | ----------------------------------------------------------------------------- | --------------------------------------- |
| `ltx-2.5-22b-distilled-transformer-bf16.safetensors`      | Distilled LTX-2.5 transformer (loaded via **UNETLoader**)                     | `ComfyUI/models/diffusion_models/`      |
| `gemma4-12b-with-proj-ltx-2.5-bf16.safetensors`           | Text encoder (Gemma 4 12B)                                                    | `ComfyUI/models/text_encoders/`         |
| `gemma4_e2b_it_bf16.safetensors`                          | Prompt enhancer (Gemma 4 E2B); from `Comfy-Org/gemma-4`, not the LTX-2.5 repo | `ComfyUI/models/text_encoders/`         |
| `ltx-2.5-video-vae-bf16.safetensors`                      | Video VAE                                                                     | `ComfyUI/models/vae/`                   |
| `ltx-2.5-audio-vae-bf16.safetensors`                      | Audio VAE                                                                     | `ComfyUI/models/vae/`                   |
| `ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors` | Spatial upscaler (2×)                                                         | `ComfyUI/models/latent_upscale_models/` |
| `ltx-2.3-22b-ic-lora-in-outpainting-0.9.safetensors`      | Inpainting/outpainting IC-LoRA (reused from LTX-2.3)                          | `ComfyUI/models/loras/`                 |

> **Note**
>
> The same IC-LoRA checkpoint is used for both inpainting and outpainting — the difference is how the mask is configured. LTX-2.5 ships them as two separate workflow files.

## How It Works

Each workflow loads a source video, defines a mask for the region to generate, and runs a two-stage pipeline. Each stage does mask-aware preprocessing and blends the generated result back into the original:

1. **Preprocess (Stage 1)** — Masks the frames (`LTXVInpaintPreprocess`), builds the empty audio-video latent at the reference video's size and length, optionally conditions on a start image, injects the IC-LoRA guide, and adds **frozen audio** (`VAEEncodeAudio` → `SetAudioRefTokens`) so the result stays consistent with the source's sound.
2. **Generate — Low Res (8 steps)** — Generates at the base resolution with the IC-LoRA conditioning; a Laplacian pyramid blend merges the output with the original content.
3. **Preprocess — Stage 2** — VAE-encodes the blended Stage-1 result, re-conditions on the high-res start image, and re-attaches the frozen audio.
4. **Generate — High Res (2 steps)** — Refines the upscaled video, again blending to keep clean mask boundaries.

> **Note**
>
> Frozen audio keeps the generated region consistent with the original sound. If you don't want that guidance — for example, when replacing a section of the frame, or to keep off-frame sound off-frame — remove the audio stream from the generation and/or adjust the conditioning.

> **Note**
>
> These workflows use `LTXAddVideoICLoRAGuideAdvanced` (not the standard `LTXAddVideoICLoRAGuide`) for mask-aware conditioning.

## Outpainting

Outpainting extends the video canvas beyond its original boundaries, generating new content in the padded region while preserving the original footage in the center.

### Step-by-Step

1. **Download and load the [outpainting workflow](https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/example_workflows/2.5/LTX-2.5_ICLoRA_Outpaint_Two_Stage_Distilled.json)** and drag it into ComfyUI. Install any missing custom nodes / models via the **Workflow Overview** panel.
2. **Load your source video** in the **LoadVideo** node.
3. **Set the target canvas size** — the source is centered in the target canvas and the surrounding pad becomes the mask (the region the model generates).
4. **Write your prompt (optional)** — with no prompt, the model extends the canvas to match the surrounding scene. With a prompt, write a regular T2V-style description of the scene; it is not an edit tool. Enable the Gemma 4 enhancer to expand a short prompt.
5. **Generate.** Runs the two-stage process above.
6. **Review and iterate** — if the boundary between original and generated content shows issues, tune the blend **dilation** (see [Customization](#customization)).

## Inpainting

Inpainting fills masked regions within the frame (removing or replacing objects) while preserving everything outside the mask.

### Step-by-Step

1. **Download and load the [inpainting workflow](https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/example_workflows/2.5/LTX-2.5_ICLoRA_Inpaint_Two_Stage_Distilled.json)** and drag it into ComfyUI. Install any missing nodes / models via the Workflow Overview panel.
2. **Load your source video** in the **LoadVideo** node.
3. **Load your mask video** — a black-and-white mask (**white = inpaint, black = keep**) that matches the source video's frame count and aspect ratio. For object replacement, a looser mask works better than a tight silhouette. Create masks with a segmentation model (e.g. SAM), manual painting, or a bounding box.
4. **Write your prompt (optional)** — with no prompt, the model fills the masked area to match the surroundings (usually removing the masked object). With a prompt, describe the **full scene** you want, not the edit (Right: "a horse walking down an empty country road, sunny afternoon"; Wrong: "replace the car with a horse").
5. **(Optional) I2V mode for replacement** — for object replacement, generate the first frame separately with the replacement composited in, then set **bypass\_i2v** off to use that frame via **LoadImage**.
6. **Generate**, then **review and iterate** — tune mask dilation and blend dilation at the boundaries (see [Customization](#customization)).

> **Warning**
>
> For **replacement**, size the mask for the *new* object, not the old one — a car-sized mask won't give the model room to generate a truck. Increase dilation or use a bounding-box mask.

## Key Nodes

* **`LTXAddVideoICLoRAGuideAdvanced`** — mask-aware version of the IC-LoRA guide node; passes mask info through so the model knows which regions to generate vs. preserve. Use the workflow defaults.
* **`LTXVInpaintPreprocess`** — prepares the masked video input for each stage (no configurable parameters).
* **`LTXVDilateVideoMask`** (inpainting) — expands the mask before processing so the model has room to blend beyond the exact edge (`spatial_radius`; `temporal_radius` default 0).
* **`LTXVLaplacianPyramidBlend`** — blends generated output with the original at mask boundaries; **dilation** (`mask_low_res_dilation`) is the key parameter for seam quality.

## Customization

### Dilation

**Mask dilation** (`LTXVDilateVideoMask`, inpainting) expands the mask itself — increase it if edges are unclean or the mask is tight. **Blend dilation** (`mask_low_res_dilation` in `LTXVLaplacianPyramidBlend`) controls how far blending extends past the mask edge — the most impactful setting for seam quality; raise it when the boundary contains low-frequency content (sky, smooth gradients).

### Prompting

Describe the **full scene**, not the region or the edit — the model uses the prompt as context for the whole frame. Outpainting prompts are optional (the model extends from context); inpainting with no prompt usually removes the masked object.

### CFG

Both stages use CFG 1. Raising it adds overhead and can oversaturate; if experimenting, stay in the 1.0–1.5 range.

## Tips & Troubleshooting

* **Prompts describe the scene, not the edit** — write a regular T2V-style prompt for the region; this isn't an editing tool.
* **Boundary seam visible** — adjust the blend `dilation` in `LTXVLaplacianPyramidBlend`.
* **Green artifacts at mask edges** — the pipeline composites green under the mask before diffusion, and traces can leak through. Try a bigger dilation, a different seed, or re-encode the mask losslessly; the blending parameter matters a lot.
* **Inpainted object doesn't match** — describe the full scene, not the edit; for complex replacements use I2V mode with a composited first frame.
* **Tight mask leaves artifacts** — increase `spatial_radius` in `LTXVDilateVideoMask` or use a looser mask.

## Technical Notes

* The IC-LoRA loader's `latent_downscale_factor` is intentionally unused — the In-Outpainting LoRA was trained at `reference_downscale_factor = 1`, so the reference is processed at the output resolution.
* Two-stage: Low Res (8 steps) → High Res (2 steps), with Laplacian pyramid blending at each stage.
* Audio is guided by frozen audio reference tokens (`VAEEncodeAudio` → `SetAudioRefTokens`) rather than generated fresh, keeping the result consistent with the source's sound.