Keep lines under twelve words. Three separate things push in that direction: captions wrap awkwardly past that length, lip-sync drifts as a take runs longer, and every extra second of dialogue is generation time you pay for. They're independent constraints, so knowing which one a long line is breaking tells you whether it actually matters.
The twelve-word rule gets passed around as a convention with no reasoning attached. The reasoning is the useful part, because two of the three reasons don't apply to every line.
Why twelve words?
Because three unrelated pressures happen to converge around the same length.
That convergence is why the number feels arbitrary and works anyway. It isn't derived from one thing — it's roughly where captions, lip-sync and cost all start to bite.
Treating it as one rule makes it feel like a restriction. Treating it as three lets you break it deliberately when only one applies.
| Constraint | What breaks | Does it apply to every line? |
|---|---|---|
| Caption wrapping | Past roughly twelve words a caption wraps to a third line, covering more of a vertical frame than the composition allows | Yes, on any line that's captioned — which is most of them, since most vertical viewers watch muted |
| Lip-sync drift | Sync degrades as a take runs longer. A short line lands; a long one loses the mouth toward the end | Only on lines delivered on camera. Voiceover and off-screen lines are unaffected |
| Clip economics | Every extra second of dialogue is generation time you pay for, and re-rolls multiply it | Yes, but it scales with how much you re-roll rather than with the line itself |
Source: Hexcoded, September 2026. Established short drama craft conventions. The twelve-word figure is a working convention rather than a measured threshold.
Where this falls short. Twelve isn't a measured boundary. It's where the format has settled, and the underlying pressures are gradual rather than sharp. A thirteen-word line isn't broken; a twenty-word line is fighting all three at once.
Which lines can be longer?
Two cases, and they're worth knowing because they buy you room where it matters.
Voiceover and off-screen dialogue. No mouth on screen means no lip-sync constraint. Captions and cost still apply, but you've removed the constraint that degrades most visibly.
Lines delivered in wide shots. Lip-sync failure is a close-up problem. At a distance where the mouth is a few pixels, drift is invisible — which is a real reason to put your longest line in your widest shot.
That second one is a blocking decision as much as a writing one, and it's the kind of thing that only becomes obvious once you've seen a long line fail in a close-up.
Put the long line in the wide shot. Lip-sync failure is a close-up problem.
How do you cut a line without losing it?
Four moves, in the order they cost least.
Cut the address. Names, titles and greetings at the front of a line are usually doing nothing the shot isn't already doing. The audience can see who's being spoken to.
Cut the hedge. "I think," "maybe," "I was wondering if." In a thirty-second episode, hedging reads as padding rather than characterisation.
Split across the cut. A long line becomes two short lines with a reverse angle between them. You've spent a shot and bought back the length, which is a real trade rather than a free one.
Move it to action. If a line is explaining something the audience could see instead, the cut is free and the scene gets better.
The order matters. The first two cost nothing. The third costs a shot. The fourth costs a rewrite and usually improves the scene, which is why it's last rather than first — it's the most valuable and the least immediate.
Where does the length actually get decided?
In the script, and in Talking Actors that's literal rather than figurative.
There's no duration setting. Video length is derived from script length and natural speaking cadence, with a live word count and a duration estimate updating as you type — six words reads immediately as roughly three seconds. The estimate is an output of the script rather than a value you configure.
Two consequences follow. Length is controlled by editing the script, and there's no setting that will compress a long script into a short runtime. And because the estimate moves as the word count moves, a take that has to hit a specific duration can be written to that target directly, with the counter and the length readout together showing how much room remains.
The ceiling is sixty seconds. Real-time pacing feedback operates up to that cap.
What does this mean for captions specifically?
Most vertical viewers watch muted, which makes captions the primary channel rather than an accessibility layer.
That reframes the writing problem. A line isn't finished when it sounds right — it's finished when it reads right in two lines of caption at phone size, with a face behind it that still has room to perform.
The specific failure to watch for: a line that fits in two caption lines when spoken slowly and wraps to three when delivered at pace. The caption breaks at render, not at write, so it's discovered late unless you're checking.
How do you control where a beat lands?
Two mechanisms, and one of them is more precise than most people realise.
Emotion control in Talking Actors has two modes. Auto infers context from the script and renders matching facial expressions and vocal inflections across the whole thing, with no tagging required. Manual takes inline triggers inserted at specific words — type a forward slash at any point in the script editor.
The distinction is granularity rather than quality. Auto reads intent from the writing; Manual takes instruction at exact positions in it.
Which matters for dialogue length in a specific way. If a particular beat has to land — the pause before an answer, the shift on one word — you can place it precisely rather than lengthening the line to signal it. That's another way of buying words back.
- Under twelve words a line. Three separate constraints converge around that length
- Captions wrap, lip-sync drifts, and every second costs generation time. They're independent, so check which one you're breaking
- Voiceover and off-screen lines carry no lip-sync constraint. Spend your length there
- Put the long line in the wide shot. Lip-sync failure is a close-up problem
- Cut the address, then the hedge, then split across the cut, then move it to action. Cheapest to most valuable
- Length is controlled by editing the script. In Talking Actors there's no duration setting at all, and the cap is sixty seconds
- Most vertical viewers watch muted, so captions are the primary channel rather than an accessibility layer
- Place a beat with an inline emotion trigger rather than lengthening a line to signal it
Under twelve words. Three separate pressures converge around that length — captions wrap awkwardly past it, lip-sync drifts as a take runs longer, and every extra second of dialogue is generation time you pay for.
Sync degrades as a take runs, so a short line lands cleanly while a longer one loses the mouth toward the end. It only affects lines delivered on camera — voiceover and off-screen dialogue carry no lip-sync constraint at all.
Two cases. Voiceover and off-screen lines, because there's no mouth on screen to desync. And lines delivered in wide shots, because lip-sync failure is a close-up problem — at a distance where the mouth is a few pixels, drift is invisible.
Four moves, cheapest first. Cut the address, since the shot already shows who's being spoken to. Cut the hedge, which reads as padding at this length. Split the line across a cut, which costs a shot. Or move it to action, which costs a rewrite and usually improves the scene.
By editing the script. There's no duration setting — length is derived from script length and speaking cadence, with a live word count and estimate updating as you type. The cap is sixty seconds, so a script that would run longer has to be cut rather than compressed.
Because most vertical viewers watch muted, which makes captions the primary channel rather than an accessibility layer. A line isn't finished when it sounds right — it's finished when it reads right in two caption lines at phone size with room for the face behind it.
The script is the runtime
Talking Actors derives length from what you write, with a live word count and duration estimate as you type. Place a beat exactly with an inline emotion trigger, and audition the voice before you spend a credit.
Try Talking ActorsMore on prompt structure, syntax and per-model wording in Prompts.