xbrush logo | Blog
Docs Pricing
English 한국어
Go to App
Docs Pricing Go to App
Insight

Sound and the Final Two Seconds of an AI Ad Video — Dialogue, Lip Sync, Audio, Endings, Revisions

Byoul Oh's avatar
Byoul Oh
Oct 01, 2026
Sound and the Final Two Seconds of an AI Ad Video — Dialogue, Lip Sync, Audio, Endings, Revisions
Contents
Tip 16. Write only the words that will actually be spokenReal example — delivery outside the quotes, words insideTip 17. Measure lip sync by the clockReal example — pin the second the dialogue endsTip 18. Keep diegetic sound and music in separate slotsReal example 1 — one 🔊 SOUND line per blockReal example 2 — an audio-only timelineTip 19. Specify the final frameReal example — a last cut labeled "ENDING"Tip 20. Change one axis per revisionReal example — same content, different modelAll 20 tips on one cardFrequently Asked QuestionsHow should I write dialogue in an AI video prompt?How long should a lip-sync line be?Can I specify background music and sound effects together?How do I specify the last scene of an ad video?Where should I start when a result isn't right?

Key takeaway — A great picture is still unusable as an ad if the sound is muddy and the ending trails off. ⑯ Write only the words that will actually be spoken, in quotes ⑰ Fit lip-sync lines to the seconds (about 8–10 English words in 3 seconds) ⑱ Keep diegetic sound and music in separate slots ⑲ Specify what stays on screen in the last 1–2 seconds ⑳ Change one axis per revision. Half of 60 published prompts had dialogue; the other half wrote "no dialogue" and named what would be heard instead.

Timeline showing dialogue, diegetic sound, and music tracks and the final frame of an AI ad video

This is part 5 of our AI ad video directing series and the last of the tips. Part 4 locked what's on screen; this time it's sound, endings, and how to revise. This is the fourth group of the xbrush Academy's 20 Tips.


Tip 16. Write only the words that will actually be spoken

Description in the instructions, speech in the dialogue slot. Mix them and the description gets read aloud. Put "a confident line introducing the product" in the dialogue slot and it may well be spoken verbatim. Write the sentence to be said, and put how it's said somewhere else.

Diagram showing the spoken line alone in the dialogue slot, with tone and pace written as direction outside it
  • Quote the line exactly as it will be spoken.

  • Keep delivery outside the quotes — "low voice, unhurried."

  • No dialogue? Write "no dialogue" and name what is heard instead.

Weaker

Stronger

The model naturally explains the product's benefits.

Dialogue: "Ten minutes off your morning." — low voice, unhurried.

Real example — delivery outside the quotes, words inside

A Korean noir dialogue scene by Korean YouTuber 머니스웨거 puts the Korean lines in quotes inside English direction, and places how they're said (quietly, flat) before the quotes.

@image4 leans over the counter, the last politeness leaving his face,
says quietly in Korean: "할머니. 그거 없으면 우리 다 죽어요." ("Grandma. Without it, we're all dead.")
... She says flat in Korean: "그건 손님 사정이고." ("That's your problem.") Camera locked off.

Source: 머니스웨거 on YouTube (youtube.com)

One beauty-ad prompt on Threads separates a narration slot from a tone slot entirely: under 🎙 Narration (13–15s) "A quiet comfort from nature — slowflow" it writes, separately, Tone: female voice-over in her 30s, low and warm, slow like speaking (Threads @ssaengcho).

Scenes without dialogue aren't left blank either. A close-up emotional performance prompt specifies what is heard instead of speech: [AUDIO] No music. Hyper-realistic, wet breathing, fabric rustling, fading station noise. (Threads @moosae). Of the 60 published prompts, 31 (52%) carried speech, and most of the other 29 wrote down the sound like this.


Tip 17. Measure lip sync by the clock

Three seconds of English holds roughly eight to ten words (in Korean, about 12–15 characters). Too long a line and the model either rushes the delivery or fakes the mouth shapes. Both show. Trimming the sentence to fit the seconds is far easier than repairing it afterwards.

Reference chart of how much dialogue fits naturally into shots of different lengths
  • Write the line, read it aloud, and time it.

  • Too long? Cut the sentence. Don't ask for faster speech.

  • Move long explanations into a caption or voice-over instead of dialogue.

Weaker (3-second shot)

Stronger (3-second shot)

"Our product is an extremely convenient tool that dramatically reduces the time your morning routine takes."

"Ten minutes off your morning."

Real example — pin the second the dialogue ends

The in-flight UGC ad in the Higgsfield guide fits six lines into 30 seconds, states that all speech ends by 25 seconds, and leaves the last five seconds wordless. Lines are short, and each is finished.

Her voice: young American accent, warm and low-key, relaxed unhurried pace,
every line a complete finished sentence, no trailing off. All speech ends by 25s.

Source: Higgsfield Blog (higgsfield.ai)

A 30-second sci-fi short shared on the Naver blog Brandmite has exactly one line of dialogue: {둘 다 끝났어.} ("Both of them are done."). In action films, the line works through timing, not length.

Source: Brandmite on Naver Blog (blog.naver.com)

If you're making a lip-sync ad where an AI model speaks directly, like xbrush's Talk To You, applying this length rule at the script stage pays off most. We cover line patterns in more depth in How to Write High-Converting Scripts for Talk-to-You.


Tip 18. Keep diegetic sound and music in separate slots

Put them in one line and one of them disappears. Lump the audio together and the model usually keeps only the music. Footsteps and the clink of a glass are what make a frame feel real, and they're the first thing to go.

Audio timeline diagram directing diegetic sound and music as separate tracks by time
  • For diegetic sound, say what is heard and when — "0:03 glass set down."

  • For music, give character and level — "soft piano, under the dialogue."

  • Adding music later? Write "no music" so nothing gets invented.

Real example 1 — one 🔊 SOUND line per block

The beauty-ad prompt above gives every two-second block its own sound line and writes diegetic sound (wind, a water drop) and music (guitar, piano) side by side with a "+".

🔊 SOUND: starts in total silence — dawn forest wind fades in very softly,
🔊 SOUND: forest wind continues + a single drop of dew falling,
🔊 SOUND: guitar melody continues + the sound of one drop of serum (ASMR),

Source: Threads @ssaengcho, translated from Korean (threads.com)

Real example 2 — an audio-only timeline

A 30-second fantasy short by Threads @crome.ai keeps an AUDIO TIMELINE block separate from the picture timeline, lists the diegetic sound per segment, and spells out "no background music."

AUDIO TIMELINE:
0.0–5.2s: realistic Korean apartment ambience, smartphone taps, natural whistling
and clear Korean dialogue; no background music.
5.2–8.7s: low room tone and near-silence during the eye-contact moment, ...

Source: Threads @crome.ai (threads.com)

Sports-ad prompts often nail it down in the first line: NO BGM, FULL SFX ONLY. Music goes on in the edit; generation delivers only diegetic sound. Generating sound effects and music separately on xbrush and combining them on the timeline follows the same logic.


Tip 19. Specify the final frame

With no defined ending, the film just stops. Without an end to aim at, the model lets the last shot trail off. Pin the closing frame and the shots before it start converging on it. Any ad you plan to publish needs somewhere for the logo, the product, or the line to sit.

Comparison of an ending that trails off without a specified final frame versus a locked front-on product shot with space left for the logo
  • Say what stays on screen for the last second or two.

  • Putting type over it? Say to keep that area clear in the same breath.

  • Making a series? Keep the final frame identical across films.

Weaker

Stronger

Show the product at the end and finish.

0:13–0:15 product front on, locked off. White wall behind; keep the bottom-right quarter clear for the logo.

Real example — a last cut labeled "ENDING"

The K-pop convenience-store lookbook we've cited before marks its last cut ENDING and specifies both the camera decelerating to a stop and the person's final pose.

08.6s CUT 9 — WS→MS, 47°, eye-level slightly low, fast push-in decelerating to still.
She struts right up to camera and freezes in a composed full-figure model pose
with a faint smile — ENDING.

Source: 존코바 on Notion (notion.site)

The beauty-ad example fades the copy in at center in the last two seconds, with the logo appearing below it in the brand color. Action prompts also write an End state: for each segment — how things stand when it ends — so the next segment picks up from there (Threads @prompt_what).


Tip 20. Change one axis per revision

Fix three things at once and you'll never know which one worked. Generation varies run to run. Change several things together and you can't tell whether the improvement was your edit or luck. One axis at a time looks slower but finds the cause in two or three passes.

  • List the axes and order them — composition → light → color → sound.

  • After each change, put the result next to the previous one and compare.

  • Keep every version that improved. It's the starting point for the next film.

Real example — same content, different model

A 10-second sports-ad prompt by Threads @oiocha567 was published in two versions: the structured original for Seedance, and a narrative conversion of the same content for Grok Imagine. The scene, camera, and sound stay the same; only one axis — model and writing style — changes, so the two results can be compared side by side.

Source: Threads @oiocha567 (threads.com)

This is why part 3 said to keep block names stable. With fixed slots, "this time I'll only change the [Light] slot" becomes possible. Academy lesson L5 (A/B — one variable at a time) is exactly this exercise.


All 20 tips on one card

Checklist card summarizing the 20 AI ad video directing tips in four groups

Group

Tips

Before you order

① Length first ② One job per film ③ Placement picks the ratio ④ Half the budget for retries ⑤ References first

The shape of a prompt

⑥ One of three shapes ⑦ Timecodes add up to the length ⑧ One action per shot ⑨ Sorted, not long ⑩ Stable block names

Locking the frame

⑪ Bind people by traits ⑫ Name the camera ⑬ Light and hour in one line ⑭ Two colors ⑮ Fill, don't forbid

Sound, dialogue, ending

⑯ Only the spoken words ⑰ Fit lines to the seconds ⑱ Split sound and music ⑲ Specify the final frame ⑳ One axis per revision

Each tip's full text and practice task is in the xbrush Academy's 20 Tips. The final post in this series maps out where to find great ad video prompts like the ones quoted here.


Frequently Asked Questions

How should I write dialogue in an AI video prompt?

Quote only the words that will actually be spoken, exactly as they'll be said, and write tone and pace outside the quotes. If you put a description like "a confident line introducing the product" in the dialogue slot, the model may read that description aloud.

How long should a lip-sync line be?

The xbrush Academy suggests roughly eight to ten English words (or about 12–15 Korean characters) for a three-second shot. Longer lines get rushed or the mouth shapes slip, so shorten the sentence or move long explanations into captions or voice-over.

Can I specify background music and sound effects together?

If you put them in one line, the model usually keeps the music and drops the diegetic sound. Write diegetic sound as what is heard and when, and music as character and level, in separate slots. If you'll add music in the edit, write "no music" at the generation stage.

How do I specify the last scene of an ad video?

State what stays on screen in the last one or two seconds, with the timing — for example, "0:13–0:15 product front on, locked off, bottom-right quarter clear for the logo." Leaving space for the logo or copy in advance makes editing much easier.

Where should I start when a result isn't right?

Change one axis at a time. Go through composition, then light, then color, then sound, and compare each new version side by side with the previous one. If you change several things at once, you can't tell what made the difference.

Share article
Contents
Tip 16. Write only the words that will actually be spokenReal example — delivery outside the quotes, words insideTip 17. Measure lip sync by the clockReal example — pin the second the dialogue endsTip 18. Keep diegetic sound and music in separate slotsReal example 1 — one 🔊 SOUND line per blockReal example 2 — an audio-only timelineTip 19. Specify the final frameReal example — a last cut labeled "ENDING"Tip 20. Change one axis per revisionReal example — same content, different modelAll 20 tips on one cardFrequently Asked QuestionsHow should I write dialogue in an AI video prompt?How long should a lip-sync line be?Can I specify background music and sound effects together?How do I specify the last scene of an ad video?Where should I start when a result isn't right?
xbrush logo
Lightweight Inc.
CEO Yunho Yeon | Business Registration 208-87-02239
E-commerce Registration 2026-Seoul Seocho-1518
Unit 306, Seoul AI Hub, 47 Maeheon-ro 8-gil, Seocho-gu, Seoul, South Korea
contact@lightweight.kr
Resources
Blog User Guide
Terms and Policy
Terms of Service Privacy Policy Cookie Policy
Customer Service
Mon–Fri 10:00 AM – 6:00 PM (KST)
+82-507-1336-9329
contact@lightweight.kr
Copyright ⓒ 2026 Lightweight Inc. All Rights Reserved.