Extension chains directly from a clip you already rendered, appending footage across sequential rounds while maintaining visual continuity. Because the model is continuing from real footage rather than a description, space, lighting and identity carry across without being restated. What drifts is anything the source doesn't show — and what compounds, round after round, is any descriptor you change between prompts.
Every continuity problem in AI video comes from generating a shot with no knowledge of the shot before it. Extension is the one operation that doesn't have that problem.
What makes extension different from generating a new shot?
The model can see where it's starting from.
A fresh generation starts from noise and your description. It has no idea what the previous shot looked like, which is why space, light and identity all have to be restated and can all still drift.
An extension chains from a previously rendered clip across sequential rounds, maintaining visual continuity. The room is already there. The light is already established. The face is already on screen. None of that has to be described, because it's visible.
That's a structurally different operation, and it's why extension holds continuity that a new generation can't.
A new shot is a guess. An extension is a continuation.
Seven models support Extend: Seedance 2.5, Seedance 2.0, Seedance 2 Fast, Seedance 2 Mini, Grok Imagine 1.5, Veo 3.1 and Veo 3.1 Fast. The model dropdown filters automatically, surfacing only the engines certified for the mode you're in.
How much can you actually append?
It depends on the model, and the two ends of the range are further apart than you'd expect.
Seedance 2.5 extends a clip that already reaches thirty seconds in a single pass, so extension there is about going beyond a length that's already long.
Veo 3.1 is the opposite case and the clearer illustration. It's locked to eight seconds per generation in Create mode, and Extend appends roughly seven seconds per run for up to around twenty runs — so a clip reaches well past two minutes, with output matching the source. For Veo, extension isn't an add-on. It's the only route to any real duration.
What holds, and what drifts?
The rule is simple: what's visible in the source holds. What isn't, drifts.
| Element | Behaviour across the join | Why |
|---|---|---|
| Framing and space | Holds well | The geometry is visible in the source |
| Aspect ratio and resolution | Inherited, not chosen | In Extend the aspect ratio comes from the source and the selector is hidden entirely |
| Lighting | Holds well | Direction, quality and colour are all present in the image |
| Character identity | Holds while the face stays in frame | The face is visible, so it's being continued rather than re-imagined |
| Wardrobe and props in shot | Holds well | Present in the source |
| Anything out of frame | Drifts | The model has no information about what it can't see |
| New motion or action | Least reliable | Requires the model to invent rather than continue |
Source: Hexcoded Creative Studio: Video documentation and picker, checked September 2026.
Where this falls short. The last two rows are where extensions go wrong. A character who turns and reveals part of the room the source never showed is asking the model to invent, and invention across a join is exactly where the seam becomes visible.
What actually causes drift across rounds?
Hexcoded's own model documentation names it, which is more than most guidance on this offers.
The mechanism is latent re-encoding compounding small variations when subject descriptors are altered in chained Extend prompts. Each round re-encodes what came before, so a descriptor you changed — even slightly, even unintentionally — doesn't just affect that round. It gets carried forward and amplified.
The stated fix is the one that follows from it: restate exact clothing, colour and subject descriptors identically in every Extend prompt, and avoid introducing conflicting ones.
Which means extension rewards copy-paste discipline over rewriting. The prompt for round four should be the prompt for round three with only the new action changed.
Restate the descriptors identically. Change only the action. Drift across rounds is almost always something you introduced.
When does the join become visible?
Three cases, and all three are about change rather than duration.
A change of motion direction. Someone walking left who begins walking right across the join. The model has momentum in the source frames and reversing it is where the physics wobble.
A change of performance intensity. A calm read becoming an angry one across the join. Emotion carries in micro-expression, and that's the thing least well described by the last few frames.
A reveal. Anything entering frame, or the camera moving to show new space. The model is generating that from nothing, in the middle of a shot that was otherwise continuous.
What these share is that the source doesn't contain the information the extension needs. Extension is excellent at continuing and unreliable at changing.
How does this change how you plan a shot?
Put the change before the join, not across it.
If a beat needs a shift — the turn, the reaction, the moment the tone changes — generate that inside a single pass and extend after it, not through it. The extension then continues a state rather than producing a transition.
The planning consequence is that you're deciding where joins fall at script stage rather than discovering them in production. Framing too: aspect ratio and resolution are inherited from the source, so both are decided in the first generation and neither can be changed later.
What carries over between rounds besides the picture?
One thing worth knowing, because it's easy to miss.
The reference pool is maintained across Create, Edit, Extend and Motion within your active session rather than resetting after each generation. So the actor, location and style references you attached to the original shot are still attached when you extend it.
That's useful and it's also a trap. If you've moved on to a different scene in the same session, references from the previous one may still be in the pool. Worth checking what's attached before you extend rather than after.
How should you actually use it?
Five steps.
Decide where the joins fall before you generate
Put them at moments of continuation, not change. A join across a turn, a reveal or a tonal shift is the one that shows.
Set framing and resolution in the first pass
Both are inherited in Extend and the selector is hidden. Whatever you choose for the opening shot is what every round after it gets.
Generate the beat that carries the change in one pass
Whatever the shot's difficult moment is, keep it inside a single generation. Extend before it or after it, not through it.
Copy the prompt forward, changing only the action
Restate clothing, colour and subject descriptors identically. Rewriting them is what compounds drift across rounds, and the model documentation says so explicitly.
Review at the join first
Scrub the two seconds either side before watching the whole clip. Seam defects are easiest to miss in a full playthrough and easiest to catch in isolation.
- Extension chains from real footage rather than a description, so space, light and identity hold for free
- What's visible in the source holds. What isn't, drifts
- Seven models support Extend. Veo 3.1 appends roughly seven seconds a round, up to about twenty rounds
- Aspect ratio and resolution are inherited, and the selector is hidden. Both are decided in the first pass
- Drift across rounds comes from changing descriptors. Copy the prompt forward and change only the action
- New motion, new action and reveals are the least reliable things to ask for across a join
- Put the change before the join, not across it. Extension continues well and changes badly
- The reference pool persists across the session. Check what's attached before you extend
Chaining directly from a clip you already rendered, appending footage across sequential rounds while maintaining visual continuity. Because the model continues from footage rather than a description, space, lighting and identity carry across without being restated.
Seven: Seedance 2.5, Seedance 2.0, Seedance 2 Fast, Seedance 2 Mini, Grok Imagine 1.5, Veo 3.1 and Veo 3.1 Fast. The model dropdown filters automatically to whichever models support your current mode.
It depends on the model. Veo 3.1 appends roughly seven seconds per run for up to around twenty runs, with output matching the source clip — so a clip reaches well past two minutes. For Veo, which is locked to eight seconds in Create mode, extension is the only route to real duration.
Because you changed a descriptor. Hexcoded's model documentation names the mechanism: latent re-encoding compounds small variations when subject descriptors are altered in chained Extend prompts. The fix is restating clothing, colour and subject descriptors identically every round and changing only the action.
No. Both are inherited from the source clip in Extend mode, and the aspect ratio selector is hidden entirely. Whatever you set in the first generation is what every subsequent round gets.
A single pass is cleaner when it can deliver what you need, because there are no joins at all. Extension earns its place when the beat needs more than a single pass allows, or when you discover you need more after seeing the take.
Continue the take, don't restart it
Extend and Edit sit alongside Create in Creative Studio, on 30+ models and one credit balance. Your reference pool persists across all four modes in a session, so a shot that needs more doesn't mean rebuilding the scene.
Open Creative StudioMore on model capability, access and rights in Models.