Key takeaway — A great picture is still unusable as an ad if the sound is muddy and the ending trails off. ⑯ Write only the words that will actually be spoken, in quotes ⑰ Fit lip-sync lines to the seconds (about 8–10 English words in 3 seconds) ⑱ Keep diegetic sound and music in separate slots ⑲ Specify what stays on screen in the last 1–2 seconds ⑳ Change one axis per revision. Half of 60 published prompts had dialogue; the other half wrote "no dialogue" and named what would be heard instead.
This is part 5 of our AI ad video directing series and the last of the tips. Part 4 locked what's on screen; this time it's sound, endings, and how to revise. This is the fourth group of the xbrush Academy's 20 Tips.
Tip 16. Write only the words that will actually be spoken
Description in the instructions, speech in the dialogue slot. Mix them and the description gets read aloud. Put "a confident line introducing the product" in the dialogue slot and it may well be spoken verbatim. Write the sentence to be said, and put how it's said somewhere else.
Quote the line exactly as it will be spoken.
Keep delivery outside the quotes — "low voice, unhurried."
No dialogue? Write "no dialogue" and name what is heard instead.
Weaker | Stronger |
|---|---|
The model naturally explains the product's benefits. | Dialogue: "Ten minutes off your morning." — low voice, unhurried. |
Real example — delivery outside the quotes, words inside
A Korean noir dialogue scene by Korean YouTuber 머니스웨거 puts the Korean lines in quotes inside English direction, and places how they're said (quietly, flat) before the quotes.
@image4 leans over the counter, the last politeness leaving his face,
says quietly in Korean: "할머니. 그거 없으면 우리 다 죽어요." ("Grandma. Without it, we're all dead.")
... She says flat in Korean: "그건 손님 사정이고." ("That's your problem.") Camera locked off.Source: 머니스웨거 on YouTube (youtube.com)
One beauty-ad prompt on Threads separates a narration slot from a tone slot entirely: under 🎙 Narration (13–15s) "A quiet comfort from nature — slowflow" it writes, separately, Tone: female voice-over in her 30s, low and warm, slow like speaking (Threads @ssaengcho).
Scenes without dialogue aren't left blank either. A close-up emotional performance prompt specifies what is heard instead of speech: [AUDIO] No music. Hyper-realistic, wet breathing, fabric rustling, fading station noise. (Threads @moosae). Of the 60 published prompts, 31 (52%) carried speech, and most of the other 29 wrote down the sound like this.
Tip 17. Measure lip sync by the clock
Three seconds of English holds roughly eight to ten words (in Korean, about 12–15 characters). Too long a line and the model either rushes the delivery or fakes the mouth shapes. Both show. Trimming the sentence to fit the seconds is far easier than repairing it afterwards.
Write the line, read it aloud, and time it.
Too long? Cut the sentence. Don't ask for faster speech.
Move long explanations into a caption or voice-over instead of dialogue.
Weaker (3-second shot) | Stronger (3-second shot) |
|---|---|
"Our product is an extremely convenient tool that dramatically reduces the time your morning routine takes." | "Ten minutes off your morning." |
Real example — pin the second the dialogue ends
The in-flight UGC ad in the Higgsfield guide fits six lines into 30 seconds, states that all speech ends by 25 seconds, and leaves the last five seconds wordless. Lines are short, and each is finished.
Her voice: young American accent, warm and low-key, relaxed unhurried pace,
every line a complete finished sentence, no trailing off. All speech ends by 25s.Source: Higgsfield Blog (higgsfield.ai)
A 30-second sci-fi short shared on the Naver blog Brandmite has exactly one line of dialogue: {둘 다 끝났어.} ("Both of them are done."). In action films, the line works through timing, not length.
Source: Brandmite on Naver Blog (blog.naver.com)
If you're making a lip-sync ad where an AI model speaks directly, like xbrush's Talk To You, applying this length rule at the script stage pays off most. We cover line patterns in more depth in How to Write High-Converting Scripts for Talk-to-You.
Tip 18. Keep diegetic sound and music in separate slots
Put them in one line and one of them disappears. Lump the audio together and the model usually keeps only the music. Footsteps and the clink of a glass are what make a frame feel real, and they're the first thing to go.
For diegetic sound, say what is heard and when — "0:03 glass set down."
For music, give character and level — "soft piano, under the dialogue."
Adding music later? Write "no music" so nothing gets invented.
Real example 1 — one 🔊 SOUND line per block
The beauty-ad prompt above gives every two-second block its own sound line and writes diegetic sound (wind, a water drop) and music (guitar, piano) side by side with a "+".
🔊 SOUND: starts in total silence — dawn forest wind fades in very softly,
🔊 SOUND: forest wind continues + a single drop of dew falling,
🔊 SOUND: guitar melody continues + the sound of one drop of serum (ASMR),Source: Threads @ssaengcho, translated from Korean (threads.com)
Real example 2 — an audio-only timeline
A 30-second fantasy short by Threads @crome.ai keeps an AUDIO TIMELINE block separate from the picture timeline, lists the diegetic sound per segment, and spells out "no background music."
AUDIO TIMELINE:
0.0–5.2s: realistic Korean apartment ambience, smartphone taps, natural whistling
and clear Korean dialogue; no background music.
5.2–8.7s: low room tone and near-silence during the eye-contact moment, ...Source: Threads @crome.ai (threads.com)
Sports-ad prompts often nail it down in the first line: NO BGM, FULL SFX ONLY. Music goes on in the edit; generation delivers only diegetic sound. Generating sound effects and music separately on xbrush and combining them on the timeline follows the same logic.
Tip 19. Specify the final frame
With no defined ending, the film just stops. Without an end to aim at, the model lets the last shot trail off. Pin the closing frame and the shots before it start converging on it. Any ad you plan to publish needs somewhere for the logo, the product, or the line to sit.
Say what stays on screen for the last second or two.
Putting type over it? Say to keep that area clear in the same breath.
Making a series? Keep the final frame identical across films.
Weaker | Stronger |
|---|---|
Show the product at the end and finish. | 0:13–0:15 product front on, locked off. White wall behind; keep the bottom-right quarter clear for the logo. |
Real example — a last cut labeled "ENDING"
The K-pop convenience-store lookbook we've cited before marks its last cut ENDING and specifies both the camera decelerating to a stop and the person's final pose.
08.6s CUT 9 — WS→MS, 47°, eye-level slightly low, fast push-in decelerating to still.
She struts right up to camera and freezes in a composed full-figure model pose
with a faint smile — ENDING.Source: 존코바 on Notion (notion.site)
The beauty-ad example fades the copy in at center in the last two seconds, with the logo appearing below it in the brand color. Action prompts also write an End state: for each segment — how things stand when it ends — so the next segment picks up from there (Threads @prompt_what).
Tip 20. Change one axis per revision
Fix three things at once and you'll never know which one worked. Generation varies run to run. Change several things together and you can't tell whether the improvement was your edit or luck. One axis at a time looks slower but finds the cause in two or three passes.
List the axes and order them — composition → light → color → sound.
After each change, put the result next to the previous one and compare.
Keep every version that improved. It's the starting point for the next film.
Real example — same content, different model
A 10-second sports-ad prompt by Threads @oiocha567 was published in two versions: the structured original for Seedance, and a narrative conversion of the same content for Grok Imagine. The scene, camera, and sound stay the same; only one axis — model and writing style — changes, so the two results can be compared side by side.
Source: Threads @oiocha567 (threads.com)
This is why part 3 said to keep block names stable. With fixed slots, "this time I'll only change the [Light] slot" becomes possible. Academy lesson L5 (A/B — one variable at a time) is exactly this exercise.
All 20 tips on one card
Group | Tips |
|---|---|
Before you order | ① Length first ② One job per film ③ Placement picks the ratio ④ Half the budget for retries ⑤ References first |
The shape of a prompt | ⑥ One of three shapes ⑦ Timecodes add up to the length ⑧ One action per shot ⑨ Sorted, not long ⑩ Stable block names |
Locking the frame | ⑪ Bind people by traits ⑫ Name the camera ⑬ Light and hour in one line ⑭ Two colors ⑮ Fill, don't forbid |
Sound, dialogue, ending | ⑯ Only the spoken words ⑰ Fit lines to the seconds ⑱ Split sound and music ⑲ Specify the final frame ⑳ One axis per revision |
Each tip's full text and practice task is in the xbrush Academy's 20 Tips. The final post in this series maps out where to find great ad video prompts like the ones quoted here.
Frequently Asked Questions
How should I write dialogue in an AI video prompt?
Quote only the words that will actually be spoken, exactly as they'll be said, and write tone and pace outside the quotes. If you put a description like "a confident line introducing the product" in the dialogue slot, the model may read that description aloud.
How long should a lip-sync line be?
The xbrush Academy suggests roughly eight to ten English words (or about 12–15 Korean characters) for a three-second shot. Longer lines get rushed or the mouth shapes slip, so shorten the sentence or move long explanations into captions or voice-over.
Can I specify background music and sound effects together?
If you put them in one line, the model usually keeps the music and drops the diegetic sound. Write diegetic sound as what is heard and when, and music as character and level, in separate slots. If you'll add music in the edit, write "no music" at the generation stage.
How do I specify the last scene of an ad video?
State what stays on screen in the last one or two seconds, with the timing — for example, "0:13–0:15 product front on, locked off, bottom-right quarter clear for the logo." Leaving space for the logo or copy in advance makes editing much easier.
Where should I start when a result isn't right?
Change one axis at a time. Go through composition, then light, then color, then sound, and compare each new version side by side with the previous one. If you change several things at once, you can't tell what made the difference.