Models

What native audio actually gets you, and where it fails

Model Seedance 2.5Model Wan 2.7Format Short dramaMarket Global English
Short answer

Native audio means the model generates sound and picture in the same pass rather than you syncing them afterwards. What it removes is a match problem — the performance and the voice came from one generation, so they agree by construction. What it costs is control: the sound is whatever the model produced, and replacing it reintroduces the sync problem you avoided.

It's listed as a feature and it's actually a workflow decision. The question isn't whether native audio is good — it's whether the sound is a decision on your project.

What does native audio actually do?

Generates the audio in the same pass as the picture, rather than as a separate step you align afterwards.

Hexcoded's own model documentation names the mechanism. Seedance 2.5 uses an Audio-Visual Diffusion Transformer architecture, processing visual frames and acoustic signals within the same generation pass — where traditional video generation creates silent footage first and overlays external audio tracks afterward. Visual movement and sound dynamics are calculated simultaneously, so footsteps, environmental ambience, surface impacts and spoken dialogue align with visible kinematics without a post-synchronisation step.

That sounds like a convenience and it's structurally more than one. When you sync separately, you're matching two things that were produced independently, and the match is something you maintain. When the model produces both together, there's nothing to match.

The clearest place this shows is dialogue in close-up, where a small sync error is immediately visible and a matched take simply doesn't have one.

What do you actually get?

Four things, and only the first is what the feature name suggests.

What you getWhy it mattersWhere it doesn't
A matched takeMouth and sound produced together, so there's no sync to maintainOnly matters where a mouth is visible. Voiceover gains nothing
Foley tied to visible actionFootsteps, surface impacts and contact sounds align with what's on screen rather than being placed against itOnly as good as the prompt. Generic acoustic cues produce generic sound
Ambience with the sceneRoom tone and incidental sound arrive matched to the pictureYou lose the ability to build the soundscape deliberately
Faster draftingOne generation instead of generate, sync, checkLess relevant once a project is locked and you're producing finals

Source: Hexcoded model documentation and Creative Studio picker, checked September 2026. Model behaviour changes with versions.

Where this falls short. The third row is the one people underestimate. Getting ambience automatically is convenient until you want a specific soundscape, at which point automatic ambience is something you're working against rather than with.

Is it one setting, or per model?

Both, and the distinction catches people.

Creative Studio presents a native sound toggle as a global control across all four video sub-modes. But per-model behaviour overrides it. Seedance 2.5 gives you a toggle you can switch off. Wan 2.7 has sound always on with no toggle at all — and the picker describes Kling 3.0 Turbo as "fast cinematic clips, audio always on."

So "does this model do native audio" has three answers rather than two: toggleable, always on, or absent. Which means the question to ask isn't whether a model supports sound — it's whether you can turn it off.

That matters more than it sounds. If you've planned to sync separately and the model generates audio you can't disable, you're mixing against a bed you didn't ask for.

The question isn't whether a model does native audio. It's whether you can turn it off.

What does it cost you?

Control, in three specific places.

Voice casting. A natively generated voice is what the model produced for that character on that take. If the voice is a decision — a specific age, accent, register — native audio doesn't give you that lever. Talking Actors is the tool where voice becomes a casting choice, because the actor carries both the face and the voice.

Consistency across a series. Hexcoded's own Seedance 2.5 documentation names the mechanism for the visual case: latent re-encoding compounds small variations when subject descriptors are altered across chained generations. The same logic applies to sound. For a recurring character across many episodes, that's the same drift problem in a different channel.

The mix. Dialogue, ambience and effects arriving as one layer means you can't balance them independently. For a finished piece that's a real limitation; for a draft it's irrelevant.

Native audio removes a match problem and adds a drift problem. Which one you'd rather have depends on how long the series runs.

That's the honest framing. For a single piece, the match problem is the bigger one and native audio wins. Across a season, consistency becomes the harder problem and separate audio starts to look better.

When should you sync separately?

Four cases, and they're all about the sound being a decision rather than an output.

When you've cast a specific voice, whether a performer or a synthetic voice you've chosen deliberately. When a character recurs across many episodes and consistency matters more than per-shot convenience. When the soundscape is doing narrative work and needs building rather than accepting. And when you're delivering to a client who may want the sound changed after the picture is approved.

That last one is the practical trap. A picture-locked edit with baked-in audio is hard to revise, and "can we change the read?" is an ordinary note.

There's a constraint worth knowing here too. Where a human creator appears, Hexcoded's terms permit light edits only — you may not re-voice a face outside the platform. So "sync separately" means using a model where audio is off, not re-voicing a delivered take externally.

If a client can request a voice change after picture lock, plan for it before you generate. Choose a model where sound can be switched off, or use Talking Actors where the voice is a casting decision. Re-voicing a licensed face outside the platform isn't permitted.

How do you use audio as an input?

The part most people miss, because it works the other way round from what you'd expect.

Audio isn't only an output. Creative Studio's video engine accepts up to 10 audio clips as references inside its 50-asset pool, and Hexcoded's own documentation describes their role as rhythmic cues, acoustic tone and synchronisation context.

Which means you can supply a track and have generation take pacing from it, rather than generating and then cutting to music. For anything rhythm-led — a montage, a beat-matched sequence, an ASMR piece — that's the difference between editing to a track and generating to one.

Where this falls short. Ten audio references is a ceiling, not a target, and the reference pool is shared. Audio references come out of the same 50 assets as your images and actors.

How should you decide, per project?

Four questions. The first yes settles it.

1

Is the voice a casting decision?

Specific age, accent, register, or matching a performer. If yes, use Talking Actors where the actor carries the voice, or a model where audio can be switched off.

2

Does the character recur across many episodes?

Audio drift behaves like visual drift, and your own model documentation names the mechanism. Across a season, consistency usually beats per-shot convenience.

3

Is the soundscape doing narrative work?

If ambience and effects carry story rather than realism, build them. Automatic ambience is something you'd be working against.

4

None of the above?

Native audio. Faster, matched by construction, and one fewer thing to maintain. Check the model lets you turn it off anyway, in case that changes.

Model capabilities described here are published specifications and observed picker behaviour current as of the publication date, not measured comparisons. Audio behaviour differs by model and changes with versions. Verify in the picker before relying on this for a delivery.

The bottom line
  • Native audio removes a match problem. Mouth and sound produced in one pass don't need syncing
  • Seedance 2.5 does it through an audio-visual architecture rather than two layers joined. That's in Hexcoded's own model documentation
  • It costs you voice casting, consistency across a series, and independent control of the mix
  • Ask whether you can turn it off, not whether the model has it. Three answers exist: toggleable, always on, absent
  • For a single piece the match problem is bigger and native audio wins. Across a season, drift is the harder problem
  • Automatic ambience is convenient until the soundscape is doing narrative work
  • Audio is an input too. Up to 10 audio references guide pacing and acoustic tone
  • Where a human creator appears you can't re-voice outside the platform. Plan the sound before you generate

Audio generated in the same pass as the picture rather than produced separately and aligned afterwards. Seedance 2.5 uses an audio-visual architecture that processes visual frames and acoustic signals together, so sound aligns with visible movement without a post-synchronisation step.

It depends whether the sound is a decision on your project. Native audio removes a match problem and adds a drift problem. For a single piece the match problem is bigger; across a season, consistency usually matters more.

On some models. Seedance 2.5 has a sound toggle you can switch off. Wan 2.7 has audio always on with no toggle, and Kling 3.0 Turbo is described in the picker as audio always on. So the useful question isn't whether a model supports sound but whether you can disable it.

Control in three places. Voice casting, since the sound is whatever the model produced. Consistency across a series, because audio can drift between generations the way visuals do. And the mix, since dialogue, ambience and effects arrive as one layer you can't balance independently.

When you've cast a specific voice, when a character recurs across many episodes, when the soundscape is doing narrative work, or when a client might request a change after picture lock. Note that where a human creator appears you can't re-voice outside the platform, so plan it before you generate.

Yes. Creative Studio's video engine accepts up to 10 audio clips as references inside its 50-asset pool, used for rhythmic cues, acoustic tone and synchronisation context. So you can generate to a track rather than editing to one afterwards.

Sound on, sound off, sound as input

Creative Studio runs 30+ models on one credit balance, with the sound toggle visible per model before you generate — and up to 10 audio references to generate against rather than cut to.

Open Creative Studio

More on model capability, access and rights in Models.