Layout to Render
Layout to Render turns a 3D viewport animation into a finished shot. The layout can be a detailed clay render or a blocky playblast of simple shapes from Blender, Unreal, or a similar tool. It guides the camera, motion, framing, and placement of objects, while an art-directed first frame and a short prompt define materials, lighting, color, and overall appearance.
The adapter is designed for camera-accurate look development and shot iteration. It keeps the camera path and object placement aligned to the layout, including the movement of simple blocking shapes. Because the result is generative, review that alignment before using it downstream.
This workflow is distinct from Motion Transfer. It uses the layout video itself as an IC-LoRA guide; it does not extract motion from one subject and apply it to another.
Supported surfaces
Prerequisites
- A current ComfyUI installation with the LTX video nodes required by the workflow, or
ltx-pipelinesand its dependencies for Python. - The LTX-2.5 distilled model stack and 2x latent spatial upscaler required by your chosen surface.
- The Layout to Render IC-LoRA.
Model files
Model card: LTX-2.5 22B IC-LoRA Layout to Render
For ComfyUI, install these version-matched files:
For Python, provide either the monolithic distilled checkpoint and Gemma model root or the equivalent split model components supported by ltx-pipelines. The native pipeline also requires the spatial upsampler and Layout to Render IC-LoRA paths.
Inputs
Both the layout video and art-directed first frame are required. The ComfyUI graph labels the first-frame input Load reference image (Look Still) and connects both inputs directly with no bypass. In Python, you set output dimensions, frame count, and frame rate explicitly.
The model card recommends a 24 fps layout clip and identifies 1920×1088 as a known-good output size. Width and height must be divisible by 64. The workflow snaps the source length down to a supported 8k+1 frame count; for example, a 72-frame clip becomes 65 frames.
Use one or two sentences that describe the finished shot and match the art-directed first frame. Do not describe the layout source — avoid words like “clay,” “3D,” or “Unreal,” which push the output toward the blockout look instead of the finished render.
Run in ComfyUI
- Load
LTX-2.5_ICLoRA_Layout_To_Render_Two_Stage_Distilled.jsonin ComfyUI. - In Source video, select the layout or clay-render clip. Its dimensions, frame rate, and valid frame count drive the generated clip.
- In Load reference image (Look Still), select the art-directed image created from the layout’s first frame.
- Enter positive and negative prompts if needed. Add an LTX API key only if you want to offload prompt encoding.
- Confirm that the required model files are installed. Keep the model and generation settings supplied with the workflow unless the model card documents a supported alternative.
- Choose a seed and review the decode tile and overlap settings for the available VRAM.
- Run the workflow and review the saved video for composition, camera motion, geometry drift, and temporal consistency.
Run with Python
Use the same references with the native IC-LoRA pipeline: pass the layout video as video conditioning and the art-directed first frame as an image at temporal position -1. The --stage-2-ic-lora option keeps the Layout to Render conditioning active during the high-resolution stage.
This example uses the model card’s known-good output size of 1920×1088. Keep width and height divisible by 64. Repeat --image PATH FRAME_INDEX STRENGTH to add mid-shot or final keyframes; the normal recipe begins with the art-directed first frame at position -1.
The example uses the monolithic-checkpoint form. If your installation uses split model files, replace the distilled checkpoint and Gemma arguments with the split-model arguments supported by your installed version.
How the conditioning works
On both surfaces, the layout clip is VAE-encoded and attached as the IC-LoRA guide. The art-directed image is encoded separately at temporal position -1, establishing the first frame’s visual treatment. Additional keyframes can reinforce that treatment later in the shot.
The ComfyUI graph uses a two-stage distilled process:
- A low-resolution composition pass uses the Layout to Render IC-LoRA and both references.
- The graph removes the guide tokens, performs a 2x latent spatial upscale, and starts a high-resolution refinement pass.
- Both references are encoded again at the new resolution and reapplied before the second pass.
- The second-stage guide tokens are removed before the final video is decoded.
Guide positions are recorded against the latent resolution where they were added. Stage 1 removes its guides before the 2× latent upscale. Stage 2 encodes and attaches both references again at the new resolution, then removes those guides before decode.
ComfyUI decode memory
Use the tiled VAE decode controls in the workflow to fit the available VRAM. Fewer, larger tiles decode faster but need more VRAM; smaller tiles lower peak memory. Keep the values supplied by the workflow unless you need to tune decode memory for your hardware.
This tiling manages VAE decode memory only. It is not the separate Native Resolution generation workflow.
Output and audio
The supplied workflow uses a lower-resolution composition pass, then a 2× latent upscale and a shorter full-resolution refinement pass. The source video’s frame rate is retained, subject to the supported frame-count rounding described above.
The ComfyUI graph creates a silent audio latent and does not pass through or mux the source soundtrack. The native pipeline generates audio with its output video, but it does not preserve the source soundtrack. Add or replace audio in a downstream editor when you need source-synchronous sound.
Limitations
- Layout to Render is designed to preserve camera motion and object placement, but it remains generative. Review alignment, geometry, and temporal consistency before using the result downstream.
- Both the layout video and art-directed first frame are required.
- Neither surface preserves source audio. ComfyUI produces silent video; Python generates new audio with the result.