Prompting Guide
To get the most out of LTX models, a strong prompt makes all the difference. The key is painting a complete picture of the story you’re telling that flows naturally from beginning to end and covers all the elements the model needs to bring your vision to life.
If you’re new to writing prompts for video generation, this guide will help you construct effective, production-ready prompts.
Table of Contents
- Key Elements to Include
- Structuring Your Prompt
- Multi-Shot Prompts
- Using the Prompt Enhancer
- Keep in Mind
- Prompting for Specific Capabilities
- Sample Prompts
- Additional Helpful Terms
Key Elements to Include
When writing a prompt, aim to include the following elements:
1. Establish the Shot
Use cinematography terms that match your intended genre. Include shot scale or category-specific characteristics to refine the visual style.
2. Set the Scene
Describe lighting conditions, color palette, surface textures, and atmosphere to establish mood and tone.
3. Describe the Action
Write the core action as a natural sequence that flows from beginning to end. Give each sentence a verb that does something — walks, turns, exhales, reaches — rather than only describing how things look. Appearance alone gives the model little to animate; describing what happens moment to moment gives it motion to follow.
4. Define the Character(s)
Include age, hairstyle, clothing, and distinguishing features. Express emotion through physical cues, not abstract labels.
5. Identify Camera Movement(s)
Specify how and when the camera moves. Describing how subjects appear after the movement helps the model complete the motion accurately.
6. Describe the Audio
Clearly describe ambient sound, music, speech, or singing.
- Place spoken dialogue in quotation marks
- Specify language and accent if needed
Structuring Your Prompt
Prompts range from a single continuous take to longer, screenplay-style scenes. Match the structure to what you’re describing rather than forcing every prompt into one shape. LTX responds best to cinematic, single-subject scenes with clear camera language, consistent lighting, and well-described audio.
Coming from another model? Don’t paste a prompt written for another video model (e.g. Kling or Seedance) into LTX unchanged — the content usually carries over, but tag syntax and shot-list formatting don’t, and tend to underperform. Rewrite it into LTX’s flowing-paragraph structure, or run the original through the prompt enhancer.
A few principles apply to any prompt:
- Keep the scene focused — a few clear characters and actions read better than a crowded frame.
- Keep lighting consistent — use one coherent light logic per shot; mixed light sources confuse the result.
- Start simple and layer — begin with the core shot, then add detail as you iterate.
Simple / Single-Shot
For a single continuous take, a short flowing description works best:
- Write your prompt as a single flowing paragraph
- Use present tense verbs for action and movement
- Match the level of detail to the shot scale (close-ups need more detail than wide shots)
- Describe camera movement relative to the subject
- Aim for roughly 4–8 descriptive sentences
Longer / Screenplay-Style
When a scene involves dialogue, multiple beats, or precise timing, write it in a screenplay style, with scene headers, character cues, and quoted dialogue, as the Sample Prompts below do. Keep the same fundamentals: present tense, physical emotion cues, and dialogue in quotation marks.
Length
Match length to complexity. A simple single shot is often 4–8 sentences; a longer screenplay-style scene can run longer, provided every sentence adds concrete visual or audio detail.
Pace the action in the prompt itself. LTX-2.5’s optional duration predictor sizes the clip to the action you describe and times it as written. It won’t stretch a moment or add a pause you didn’t prompt for. Write the beats you want into the prompt (“she pauses”, “a beat of silence”) so they’re part of the action, or set an explicit duration to give the whole sequence more room.
Multi-Shot Prompts
The guidance above describes a single continuous shot (one camera take). LTX-2.5 can also generate multi-shot scenes: several distinct shots joined by explicit cuts inside one prompt.
Write the full scene as one chronological paragraph (or a short sequence of sentences). Do not use a shot list, numbered beats, or screenplay sluglines unless you also describe the cut in prose.
How Multi-Shot Differs from Single-Shot
What to Include at Every Cut
- Name the transition in natural language — e.g. “A hard cut transitions to…”, “The view cuts to a close-up of…”, “A match cut connects…”, “The image dissolves into…”.
- Re-establish the new shot — shot scale, camera angle, who or what is in frame, and lighting if it changed.
- Keep identity consistent — reuse the same visual identifiers for recurring people or objects (“the woman in the red coat, earlier at the table, now…”).
- State audio continuity — e.g. “the piano score continues across the cut” or “the dialogue drops; only wind remains.”
Tips for Strong Multi-Shot Prompts
- Prefer 2–4 shots in one generation; more cuts usually need clearer, shorter beats per shot.
- Give each shot a clear job (establish → detail → reaction, or wide → medium → close-up).
- Keep action chronological e.g. “Initially…”, “A moment later…”, “Simultaneously…”.
- The same rules as single-shot apply: present tense, physical emotion cues, quoted dialogue, concrete camera language.
- Avoid conflicting geography or unexplained costume changes between cuts unless the cut is meant to jump time or place and you say so.
Multi-Shot Example
A wide shot frames a rainy city intersection at dusk, neon signs reflecting on wet asphalt. A young woman in a yellow raincoat walks down the sidewalk toward camera, carrying a small bag, while the rain falls and cars drive past behind her. Soft synth music and traffic noise fill the air. The shot transitions to a medium close-up of her face under the hood, raindrops catching the neon as she looks off-screen left; the synth score continues across the cut, traffic muffled. She speaks quietly to herself, “He’s late.” A hard cut jumps to a low-angle shot of a man’s scuffed boots stepping into a puddle at the curb; the music drops to a low drone. The man she has been waiting for — short dark hair, soaked jacket — lifts his head into frame as he smiles at her off-screen. A bus rumbles past them.
When to Stay Single-Shot
Use a single continuous take when you want unbroken camera motion, intimate performance, or dialogue that must stay lip-synced in one framing. For image-to-video from a first frame, prefer a single continuous take unless you intentionally describe a cut away from that opening image.
Using the Prompt Enhancer
This section applies to the local open-source LTX-2.5 paths: native ltx-pipelines and the official ComfyUI templates.
For direct API requests, do not use --enhance-prompt.
LTX pipelines include an optional prompt enhancer: an LLM rewrite pass (the Gemma 4 E2B model) that expands your prompt using a system prompt tuned for text-to-video or image-to-video before generation runs. It’s enabled with --enhance-prompt in ltx-pipelines and via a dedicated node in the official ComfyUI templates, and it can be turned off in both.
We recommend using it when your prompt is short, rough, or was originally written for a different model. It helps least when your prompt already follows the structure detailed in this post. The enhancer runs an extra inference pass, so it adds some latency; leave it off if you’d rather submit your prompt exactly as written.
Keep in Mind
A couple of things the model still handles unevenly:
- On-screen text — LTX-2.5 improves short-text accuracy and preserves fine details better than earlier versions, but exact spelling and consistency across frames are not guaranteed. Keep text short and prominent, verify it throughout the clip, and add critical titles, labels, or logos in post.
- Complex physics — highly chaotic motion can introduce artifacts; simpler, more plausible motion is more reliable.
Prompting for Specific Capabilities
Some LTX features have their own prompting patterns.
Dub-It (Speech Replacement)
The Dub-It IC-LoRA is a video-to-video tool that replaces spoken dialogue in existing video. Unlike text-to-video generation, you provide a source video and write a prompt describing what the speaker should say instead.
Dub-It can be used for dubbing into other languages or for rephrasing dialogue in the original language.
Languages currently validated: English, French, Spanish, German, Russian.
Prompt Template
Example:
You can add details about emotion or delivery style to the prompt.
Requirements
- Provide the full dialogue text — the model follows the content of the prompt. It does not translate dialogue for you.
- Use native script — write dialogue in the alphabet of the target language (e.g., Cyrillic for Russian, Chinese characters for Mandarin).
- Single speaker — the beta IC-LoRA does not distinguish between multiple speakers.
Best Practices
- Match audio length — keep your prompt at roughly the same timing and syllable length as the original dialogue. Slightly longer is better than too short.
- Prompt too long: the model might skip words.
- Prompt too short: the output might sound slow and unnatural.
Sample Prompts
Example 1
Prompt:
EXT. TOWN STREET – MORNING – LIVE NEWS BROADCAST The shot opens on a news reporter standing in front of a row of cordoned-off cars, yellow caution tape fluttering behind him. The light is warm, early sun catching the camera lens. A faint hum of chatter and distant drilling fills the air. The reporter, composed but smiling nervously, looks directly into the camera, microphone in hand. Reporter: “Thank you, Sylvia. And yes — this is a sentence I never thought I’d say on live television — but this morning, here in the quiet town of New Castle, Vermont… black gold has been found!” He gestures toward the field behind him. “If my cameraman can pan over, you’ll see what all the excitement’s about.” The camera pans right, slowly revealing a construction site surrounded by workers in hard hats. A beat of silence — then, with a sudden roar, a geyser of oil erupts from the ground and blasts upward in a violent plume. Workers cheer and scramble as the black stream glistens in the morning light. Reporter (off-screen, shouting over the noise): “There it is, folks — a moment New Castle will never forget!” The camera catches sunlight gleaming off the oil mist, then pulls back to reveal the whole scene: a small town silhouetted against the wild fountain of oil.
Example 2
Prompt:
A wide shot opens in a warm, sunlit frog yoga studio with a tactile, felt-and-fabric look. Golden morning light pours through tall wooden-framed windows, lush green foliage outside, thin wisps of incense smoke curling through the air; potted plants and wooden shelves line the softly blurred background. One housefly flies lazily around the room, flitting in and out of the smoke. A large green frog instructor sits in lotus position at the center on a woven straw mat, wearing an orange robe, eyes gently closed, a serene half-smile, hands resting on his knees. Behind him, rows of smaller green frogs sit on their own woven mats, throats swelling as they chant a deep, resonant meditative “Om” in unison, a low male vocal drone with rich harmonic overtones, no instruments, no music. Soft pond ambience underneath, and the faint buzz of the fly.
The instructor breathes in slowly, then speaks in a deep, calm voice, drawing out each word. “We are one… with the pond.” The frogs answer, chanting in unison: “Om…” He smiles faintly. “We are one… with the mud.” Again the frogs chant together, “Om…” a slow beat. “We are one… with the flies.” A pause.
The camera pans slowly left to a small frog in the front row, who twitches, eyes darting as the fly drifts past. Suddenly its tongue snaps out, catching the buzzing housefly mid-air and pulling it back into its mouth.
The master exhales slowly, still serene, eyes still closed. “But we do not chase the flies…” Beat. “…not during class.”
The guilty frog lowers its head, folding its hands back into a meditative pose as the others resume their deep, resonant chant, throats swelling: “Om…” A lingering shot holds on the guilty frog.
Additional Helpful Terms
This list is not exhaustive, but provides useful examples for shaping your results.
Categories
Animation — Stop-motion · 2D / 3D animation · Claymation · Hand-drawn
Stylized — Comic book · Cyberpunk · 8-bit pixel · Surreal · Minimalist · Painterly · Illustrated
Cinematic — Period drama · Film noir · Fantasy · Epic space opera · Thriller · Modern romance · Experimental film · Arthouse · Documentary
Visual Details
Lighting — Flickering candles · Neon glow · Natural sunlight · Dramatic shadows
Textures — Rough stone · Smooth metal · Worn fabric · Glossy surfaces
Color Palette — Vibrant · Muted · Monochromatic · High contrast
Atmosphere — Fog · Rain · Dust · Smoke · Particles
Sound and Voice
Ambient Settings — Coffeeshop noise · Wind and rain · Forest ambience with birds
Dialogue Style — Energetic announcer · Resonant voice with gravitas · Distorted radio-style · Robotic monotone · Childlike curiosity
Volume — Whisper · Mutter · Shout · Scream
Technical Style Markers
Camera Language — Follows · Tracks · Pans across · Circles around · Tilts upward · Pushes in / pulls back · Overhead view · Handheld movement · Over-the-shoulder · Wide establishing shot · Static frame
Film Characteristics — Film grain · Lens flares · Pixelated edges · Jittery stop-motion
Scale Indicators — Expansive · Epic · Intimate · Claustrophobic
Pacing & Temporal Effects — Slow motion · Time-lapse · Rapid cuts · Lingering shot · Continuous shot · Freeze-frame · Fade-in / fade-out · Seamless transition · Sudden stop
Visual Effects — Particle systems · Motion blur · Depth of field