Duration decides it more often than anything else. Seedance 2.5 holds up to thirty seconds in one pass, so a long unbroken read goes there. Veo 3.1 is locked to eight seconds per generation, which fits most single lines and chains through Extend for more. Kling 3.0 Omni runs three to fifteen seconds. If the scene is a face delivering a script rather than a scene at all, none of the three is the right tool.
Most model comparisons rank quality. For dialogue that's the wrong axis, because the constraint that bites first is how long a line takes to say.
Why is duration the deciding factor?
Because dialogue has a fixed length and generation doesn't.
A line takes as long as it takes to say. You can shorten the writing, but you can't speed up delivery without it reading as rushed. So the question isn't which model handles faces best — it's which model's duration window fits the line you wrote.
That reverses the usual order. Normally you pick a model and work within it. For dialogue you write the line, measure it, and pick the model that holds it in one pass.
The alternative is splitting across a cut, which is a legitimate choice and costs you a shot.
Write the line, then pick the model. Not the other way round.
What does each one actually give you?
From the picker rather than from vendor announcements, which is a distinction worth keeping.
| Model | Duration | Resolution | Audio | Best dialogue use |
|---|---|---|---|---|
| Seedance 2.5 | 4–30s in a single pass | 480 / 720 / 1080 | Toggle, on or off | A long unbroken read. The only one of the three that holds thirty seconds without a join |
| Veo 3.1 | 8s locked per generation; roughly 7s per Extend run, up to about 20 runs | 720 / 1080 / 4K | Native, generated with the picture | Most single lines. Eight seconds covers a lot of dialogue, and it's the only 4K option |
| Kling 3.0 Omni | 3–15s | 720 / 1080 | Toggle | Mid-length lines with heavy reference control |
| Kling 3.0 Turbo | Shorter, speed-optimised | 720 / 1080 | Always on, no toggle | Fast iteration while the read is still being worked out |
Source: Hexcoded Creative Studio picker and model documentation, plus Google's Veo documentation, checked September 2026. Capabilities change with model versions.
Where this falls short. This table tells you what fits, not what performs. We haven't measured lip-sync accuracy across these models and we're not going to publish an impression of it as though we had. Duration, resolution and audio behaviour are checkable facts; "which one does mouths better" isn't, without a test.
How long is a line, actually?
Worth calibrating, because the numbers above only mean something against real dialogue.
The short drama convention is under twelve words a line. At ordinary speaking pace that's roughly four to six seconds — comfortably inside every model in the table.
Which means for single lines, duration isn't the constraint at all. It becomes one when you're running an exchange without cutting, or a monologue, or a beat where someone speaks and then reacts in the same take.
Hexcoded's Talking Actors gives a useful calibration here even if you're not using it: it derives duration from script length and shows a live estimate as you type, where six words reads as roughly three seconds. That ratio transfers as a rough guide wherever you're generating.
What about the audio?
Three different behaviours across these models, and it's a real decision rather than a detail.
Seedance 2.5 and Kling 3.0 Omni both give you a toggle — sound on or off, your choice. Veo 3.1 generates audio natively with the picture. Kling 3.0 Turbo has audio always on, with no toggle at all.
For dialogue that matters in one specific way. If you've cast a voice deliberately, or you're planning to sync a separate recording, a model with always-on audio is generating a bed you'll be working against. The useful question isn't whether a model does native audio — it's whether you can turn it off.
If the voice is a casting decision, check the toggle before the resolution.
Does reference handling change the answer?
For a recurring character, yes — and the ceilings differ sharply.
Creative Studio's video engine takes up to 50 reference assets in one generation context, composed of up to 30 images, 30 actors, 10 videos and 10 audio. But the per-model ceiling varies within that. Seedance 2.5 allows the full 50. Wan 2.7, for comparison, allows 8.
For a dialogue scene in a series, the reference you care about is the actor. Save the character once, call the handle, and the model is being shown who this is rather than being told. That matters more for identity consistency than any difference between these three models.
Kling 3.0 Omni is the strongest of the three on reference control specifically — frames, video and elements — which is why it earns its place for a mid-length line where the character has to be exactly right.
When is none of these the right answer?
When the shot is a face delivering a script and nothing else is happening.
That's what Talking Actors is for, and it's narrower by design. You pick an actor, write the script, and get a video — no model choice, because the job is defined. Duration comes from the script rather than a slider, capped at sixty seconds, with a live word count as you type.
It also has something Creative Studio doesn't: emotion control. Auto infers expression and delivery from the script; Manual takes inline triggers placed at exact words. For a piece where a specific beat has to land, that's more precise than anything a prompt will get you.
The test is whether the shot is a scene or a delivery. A scene — with blocking, environment, a second presence, camera movement — is Creative Studio. A person talking to camera is Talking Actors.
How should you decide, shot by shot?
Four questions, in order. The first that gives you a hard constraint settles it.
Is this a scene, or a person talking?
A person delivering to camera is Talking Actors — the script sets the length and emotion control is more precise than a prompt. Anything with blocking, environment or camera movement is Creative Studio.
How long is the line, unbroken?
Under eight seconds, all three work. Up to fifteen, Seedance 2.5 or Kling 3.0 Omni. Up to thirty in one pass, only Seedance 2.5.
Is the voice a casting decision?
If yes, you need a model where audio can be switched off. Veo 3.1 generates it natively and Kling 3.0 Turbo has it always on.
Does the character recur?
Then attach the saved actor reference rather than describing them, whichever model you chose. That does more for consistency than the model difference does.
- Write the line first, then pick the model. Duration is the constraint that bites before anything else
- Under twelve words is roughly four to six seconds — inside every model here. Duration only binds on longer takes
- Seedance 2.5 is the only one of the three holding thirty seconds in a single pass
- Veo 3.1 is locked to eight seconds per generation, chains through Extend, and is the only 4K option
- Kling 3.0 Omni runs three to fifteen seconds with the strongest reference control of the three
- Kling 3.0 Turbo has audio always on. If the voice is cast, that's a constraint rather than a feature
- If the shot is a face delivering a script, none of these is the tool. That's Talking Actors
- Attach the saved actor reference whichever model you pick. It does more than the model choice does
It depends on the length of the line. Seedance 2.5 holds up to thirty seconds in one pass, Veo 3.1 is locked to eight seconds per generation, and Kling 3.0 Omni runs three to fifteen. Since most single lines are four to six seconds, duration only becomes the deciding factor on longer takes.
The short drama convention is under twelve words a line, which at ordinary pace is roughly four to six seconds. As a rough calibration, six words reads as about three seconds — that's the ratio Talking Actors uses when it estimates duration from script length.
On some models. Seedance 2.5 and Kling 3.0 Omni both offer a sound toggle. Veo 3.1 generates audio natively with the picture, and Kling 3.0 Turbo has audio always on with no toggle. If you've cast a voice deliberately, check the toggle before anything else.
We haven't measured it, so we're not going to tell you. What's checkable is duration, resolution, audio behaviour and reference ceilings — and for dialogue those decide the shot more often than sync quality differences would.
When the shot is a face delivering a script and nothing else is happening. Duration comes from the script rather than a slider, capped at sixty seconds, and emotion control lets you place a beat at an exact word — which is more precise than a prompt.
Yes. Model access is tiered — the entry plan carries a core set rather than the full roster, and 4K sits above the entry tiers. So which model suits a shot and which models you can reach are separate questions.
Pick per shot, not per project
Creative Studio runs 30+ models on one credit balance, with duration, resolution, aspect ratio and sound behaviour visible per model before you generate. Switching for one line costs a decision rather than a subscription.
Open Creative StudioMore on model capability, access and rights in Models.