Key takeaway — Half the reason AI ad videos wobble is that things which should have been decided before writing the prompt were left blank. ① Length as a number ② One job per film ③ Aspect ratio set by the placement ④ Half the budget reserved for retries ⑤ References first. Across 60 publicly shared AI video prompts, not one left the length unstated, and two-thirds attached reference images.
This is part 2 of our AI ad video directing series. It covers the first group of the xbrush Academy's 20 Tips for Directing That Lands: what to decide before you order. Each tip comes with a short excerpt from a real, publicly shared AI video prompt that puts it into practice.
Tip 1. Pin the length down first
Length sets the shot count, and the shot count sets everything else. Leave the length out and the model picks its own cuts. Reading 30 seconds as eight shots instead of four doubles how much action lands in each one, and that decision drags lighting, camera, and dialogue along with it.
Write the seconds as a number — "15 seconds," not "short."
Set the shot count in the same breath. One shot per 3–4 seconds is the baseline.
If the placement is already fixed, take that surface's limit as your length.
Weaker | Stronger |
|---|---|
Make me a nice coffee ad. | 15 seconds, 4 shots. Cut every 3–4 seconds. |
Real example — length and shot count in the first line
The UGC ad example in Higgsfield's official prompting guide settles the aspect ratio, length, frame rate, and shot count in its very first sentence.
Vertical 9:16 UGC video, 30 seconds total, 24fps, shot on iPhone 14 Pro,
... natural vlog pacing, 8 shots, one consistent framing with small variationsSource: Higgsfield Blog — "Seedance 2.5: Complete Prompting Guide" (higgsfield.ai)
Korean-language prompts do the same. A parody of Korean household-goods TV commercials shared on Threads opens like this:
Seedance 2.5 T2V, 30초, 16:9, 24fps. 한국 생활용품 TV 광고 문법을 빌린 고품질 실사 풍자 코미디.
(Seedance 2.5 T2V, 30 seconds, 16:9, 24fps. A high-quality live-action satire borrowing
the grammar of Korean household-goods TV ads.)Source: Threads @sorigrim.luv, "The Money-Eating Hippo" (threads.com)
Count across 60 published prompts and 32 ran 30 seconds and 15 ran 15 seconds — 78% between them. Not one left the length unstated.
Tip 2. One job per film
Product demo plus brand mood plus how-to means all three come out blurred. Fifteen seconds holds less than you think. Give it two jobs and the model returns a compromise that does half of each, which is useless for either purpose.
Write one sentence about what a viewer should do after watching, then start.
Pick the strand first — ads, lookbooks, tutorials, and music videos follow different grammars.
Anything you want in but that fights the job goes on the list for the next film.
Weaker | Stronger |
|---|---|
Show the features, carry the brand feeling, and explain how to use it. | Job: a first-time user learns how the lid opens. Mood shots go in the next film. |
Real example — a tutorial that is only a tutorial
The capsule coffee machine tutorial in a GitHub repository of ByteDance Seedance examples states a single purpose in the first sentence and spends all 30 seconds on setup steps (Step 1–6). There isn't a single brand-mood shot.
A 30-second tutorial video on installing and using a capsule coffee machine.
0-2s: the opening title card ...
2-5s, Step 1: install the water tank, reference @image1, medium shot from a slightly high angle ...Source: GitHub — AtlasCloudAI/awesome-seedance-2.5-prompts-skills (github.com)
Tip 3. Let the placement choose the aspect ratio
Build it wide, crop it tall, and you cut people in half. A frame composed for 16:9 doesn't put the subject in the middle. Crop that to 9:16 and the face rides the edge or the product leaves the frame entirely. Composition has to know its ratio, which is why the ratio belongs in the first line.
Name where it will run, then derive the ratio from that.
If you need both wide and tall, make both. Don't ask one to cover the other.
For vertical placements, build the caption's bottom margin into the composition from the start.
Weaker | Stronger |
|---|---|
I'll make it 16:9 and crop it for Reels later. | 9:16 vertical. Eyeline a third down from the top; leave the bottom 20% clear for captions. |
Real example — vertical UGC framed vertically
The Higgsfield UGC example above leads with "Vertical 9:16" and then locks a vertical selfie composition for the whole film: "handheld selfie framing from her seat for the whole video." It doesn't stop at naming the ratio — it specifies a way of shooting that fits it.
Vertical prompts were a minority among the 60 we looked at; publicly shared examples are overwhelmingly cinematic and horizontal. If you're making a Reels or Shorts ad, don't lift a horizontal example as-is — rewrite the ratio and composition lines first.
Tip 4. Reserve half the budget for the second try
It rarely lands the first time. Spend the whole budget on attempt one and there is nothing left to correct with, so you end up shipping something you don't like. Splitting the budget up front keeps your judgement free.
Check the credit estimate the agent quotes before it runs. (xbrush quotes the expected credits before any video or audio job.)
Confirm the composition at a short length first, then go up to full length.
Fix one axis at a time — change several and you won't know which one worked.
Real example — a cheap storyboard before an expensive video
A high-school pop-dance music video prompt shared on Threads came with an image prompt that generates a 12-panel storyboard sheet first. One image confirms the composition and order; the sheet you like then goes into the video as a reference.
Create a 16:9 storyboard sheet with 12 cinematic panels for a raw, fresh,
pop dance performance set in a bright school environment.Source: Threads @jejudol_ai (threads.com)
Video generation costs far more credits and waiting time than images. Settle composition at the image stage and the same budget leaves more room for video retries.
Tip 5. Have your references ready before you start
A face, a product, or a color described only in words changes every shot. "Woman in her thirties, bob cut" points at thousands of faces, and the model picks a different one each time. One image narrows that field to one.
Anything that appears in two or more shots — person, product, space — gets an image.
Say what each reference is for: the face, the outfit, or the color.
No reference? Make the first shot, then use that result as the reference for the next.
Real example 1 — say what to take from the reference
The headphone ad in the Higgsfield guide attaches a person reference and states plainly that the face and wardrobe come from it, but the reference photo's background does not.
ACTIVE REFERENCES: The girl, identity and wardrobe locked to the reference,
ignore the reference backdrop, ...Source: Higgsfield Blog (higgsfield.ai)
Real example 2 — multiple people by number and role
A four-person dance music video prompt numbers its four reference images and fixes each performer's position on screen as a role.
IDENTITY LOCK: @Image1=A, center female; @Image2=B, female; @Image3=C, male; @Image4=D, female.
... Fixed roles: A center, B left, C right, D rear/center. No new people.Source: Threads @angga.abo (threads.com)
Of the 60 published prompts, 40 (67%) attached at least one reference, and the more a person recurred, the more images were attached. One palace ballroom one-take assigned 19 references across scenes, props, and people.
Try it on xbrush
A first order that covers all five fits in a few lines. Hand it to the agent in the xbrush workspace along with a single product photo.
Job: someone seeing this tumbler for the first time learns about its one-touch lid.
Placement: Instagram Reels → 9:16 vertical, bottom 20% clear for captions.
Length: 15 seconds, 4 shots (3–4 seconds each).
Reference: attached product photo — take the shape, color, and logo only; ignore the background.
Show me shot 1 at 5 seconds first to check the composition, then go to 15 seconds if it works.To practice hands-on, L3 (one idea into four placements) and L6 (turn a photo into video) in the Academy's 10 Hands-On Steps connect directly to this group.
Next up: the shape of the prompt itself — three structures, timecodes, and one action per shot.
Frequently Asked Questions
How long should an AI ad video be?
Follow the length your placement calls for. Among 60 publicly shared AI video prompts, 32 ran 30 seconds and 15 ran 15 seconds. Whatever the length, write it as a number and set the shot count alongside it, at roughly one shot per 3–4 seconds.
Can't I make a horizontal video and crop it to vertical?
It isn't recommended. Horizontal compositions don't keep the person or product centered, so cropping to 9:16 often cuts off faces or products. If you need both horizontal and vertical, generating each separately gives better results.
Do I really need reference images?
If the same person or product appears in two or more shots, effectively yes. Described only in words, faces and products change from shot to shot. When you attach a reference, also state what to take from it — face, outfit, product, or background — so unwanted elements don't come along.
How much budget should I keep for regeneration?
The Academy recommends keeping half the budget for retries. Check the composition at a short length first, and change only one thing per revision; you'll reach the result you want faster on the same budget.
Can I combine a product intro and brand mood in one video?
In an ad of around 15 seconds, two jobs blur both. Decide in one sentence what the viewer should do after watching, and move anything that doesn't serve it into the next film. It's faster, and the result is clearer.