A talking head is the wrong format when the information is spatial, when the content is a process, when the emotional register is intimate, when the piece needs to run past about a minute, and when the speaker's identity doesn't matter. That fourth one is a hard limit rather than a judgement — Talking Actors caps at sixty seconds.
Talking heads are the easiest AI video to produce. That's exactly why they get used for material they can't carry — the format is chosen by production convenience rather than by what the content needs.
What is a talking head actually good at?
Establishing a person and delivering a position.
When it works, it's because the speaker's presence is part of the message. A founder explaining a decision. A teacher establishing authority before the demonstration. Someone taking a view and being seen to take it.
The common thread is that the face is doing work. If you could replace the speaker with a voiceover and lose nothing, the format is decoration.
If a voiceover over other footage would lose nothing, the face isn't doing work.
Five cases where it's the wrong choice
Each with what to reach for instead.
| When | Why it fails | What to use instead |
|---|---|---|
| The information is spatial | A person describing where things are, or how something is laid out, asks the viewer to build a picture from words | Show the space. A wide shot does in two seconds what thirty seconds of description does badly |
| The content is a process | Steps described are harder to follow than steps shown, and the viewer can't check their progress against the speaker | Numbered visual steps, or a demonstration |
| The register is intimate | Direct address is a public mode. Confession, vulnerability and interiority sit awkwardly delivered to camera | A narrative scene, or voiceover over observed action |
| It has to run past about a minute | Talking Actors caps at sixty seconds, and a single static frame stops holding a viewer well before that | Cut coverage in Creative Studio, or restructure as shorter pieces |
| The speaker's identity doesn't matter | If the message would land identically from anyone, the face is occupying the frame without earning it | Faceless mode — voiceover and product visuals with no presenter |
Source: Hexcoded Talking Actors documentation and established craft, September 2026. Not measured audience response.
Where this falls short. The last row has an exception worth naming. A recurring presenter builds familiarity across a series even when any individual message doesn't need them — that's a channel decision rather than a per-video one, and it's a legitimate reason to keep a face on screen.
Why does the format get over-used?
Because production difficulty and communicative fit point in opposite directions.
A talking head is one shot, one identity, one setup. No continuity to manage, no coverage to plan, no shot-to-shot matching. It's the lowest-risk thing you can generate, and everything that makes it low-risk also makes it low-information.
There's a structural nudge too. In Talking Actors, actor selection is mandatory — the primary action stays inactive and reads "pick an actor to start" until one is chosen. Casting is the gate on generation rather than the last step before it, which means the tool asks you for a face before it asks you anything else.
That's the right design for the job it does. It's worth noticing as a nudge rather than a recommendation.
What about the length ceiling?
There's a hard one, which makes this easier to plan around than most format questions.
Talking Actors caps at sixty seconds. Duration isn't a setting — it's derived from script length and natural speaking cadence, with a live word count and estimate updating as you type. There's no slider that compresses a long script into a short runtime, so a piece that has to hit a length is written to it.
Within that minute, attention holds for roughly as long as the novelty of the face lasts. For a vertical format where episodes run thirty to ninety seconds, that's rarely the binding constraint — the piece ends before the ceiling.
Longer than a minute and the format isn't this tool's job. Cut to what they're describing, change the frame size, bring in a second element, or restructure into shorter pieces. Three sixty-second pieces frequently outperform one three-minute one, and they're easier to generate.
What does faceless mode actually do?
It's the option people skip, and it's a real answer rather than a fallback.
Faceless mode renders voiceover and product visuals with no on-screen presenter. It's available in Talking Actors, in Make a Custom Video, and in URL → Ad — where it appears as "faceless video" alongside the actor filters.
Two reasons to reach for it deliberately. When the product should carry the piece rather than a presenter, which is most of the time in commerce video. And in Make a Custom Video specifically, where selecting an actor carries a credit surcharge and faceless mode doesn't — so it's cheaper as well as sometimes better.
What if the speaker's identity matters but you don't have one?
Two routes, and they have different rights positions.
Generate an actor from a text description, and nobody real is depicted. That's the low-friction option and the right default for anything where the face just needs to be a face.
Or use a human creator — a real person who accepted a licence with a liveness check and face-match passed. That's the option when the presence has to read as real to a client or a reviewer. It needs a paid tier above the entry plan, and the licence is digital-only.
Either way the actor saves to your library with a handle and works across every tool, so the decision is made once rather than per video.
How do you decide?
Three questions before you commit to the format.
Would a voiceover over other footage lose anything?
If the answer is no, the face isn't earning the frame. Use faceless mode and put the words over the footage.
Is the information spatial or procedural?
Anything about where things are or what order to do them in is better shown. Describing a layout is asking the viewer to do work you could have done.
How long does it need to run?
Under a minute, Talking Actors holds it. Over that, it can't — the cap is sixty seconds and duration comes from the script rather than a setting. Restructure, or move to a tool where you set the length.
- Talking heads work when the speaker's presence is part of the message. Otherwise the face is decoration
- If a voiceover over other footage would lose nothing, the face isn't doing work
- Spatial information should be shown. A wide shot beats thirty seconds of description
- Processes should be demonstrated. Described steps are harder to follow and the viewer can't check progress
- Intimate registers sit awkwardly in direct address. Direct address is a public mode
- Talking Actors caps at sixty seconds, and duration comes from the script rather than a slider. Past that, use a different tool
- Faceless mode is a real answer, not a fallback. In Make a Custom Video it also avoids the actor surcharge
- A recurring presenter is a channel decision. That's a legitimate reason to keep a face on screen
When the information is spatial, when the content is a process, when the emotional register is intimate, when it needs to run past about a minute, or when the speaker's identity doesn't matter to the message. In most of those, showing beats describing.
In Talking Actors, sixty seconds. Duration is derived from script length and speaking cadence rather than set directly, with a live estimate as you type — so length is controlled by editing the script. There's no setting that compresses a long script into a short runtime.
Establishing a person and delivering a position — a founder explaining a decision, a teacher establishing authority, anyone taking a view and being seen to take it. The test is whether the face is doing work a voiceover couldn't.
Because it's the easiest thing to generate. One shot, one identity, one setup, no continuity to manage and no shot-matching. Everything that makes it low-risk also makes it low-information, so it gets chosen by production convenience rather than fit.
Numbered visual steps or a demonstration. Steps described are harder to follow than steps shown, and a viewer watching a demonstration can check their own progress against it, which they can't do against a speaker.
Voiceover and product visuals with no on-screen presenter, available in Talking Actors, Make a Custom Video and URL → Ad. It's the right choice when the product should carry the piece — and in Make a Custom Video it also avoids the credit surcharge that selecting an actor adds.
The format the content needs
Talking Actors when the face carries the message, with the script setting the length. Faceless mode when the product should. Make a Custom Video for scenes that need more than one shot. Same actor library, same credit balance.
See the toolsMore on tools, features and how the platform fits together in Product.