Talking Actors has two emotion modes. Auto infers context from the script and renders matching expression across the whole clip, with no tagging. Manual takes inline triggers — a forward slash at any point in the script editor — which places a beat at a specific word rather than describing a mood. For a clip written against a sixty-second cap, that's how you land a pause without spending words signalling it.
In every other prompt field, emotion is an adjective applied to a whole generation. Here it's a marker you position between two words, and the difference matters more than it sounds.
Why an adjective isn't precise enough
Because it colours everything equally.
Write "she speaks nervously" into a prompt and the nervousness applies across the clip. Which is fine when the whole delivery is nervous, and useless when what you actually want is one beat of hesitation before a specific word.
That's a common need in dialogue. The pause before an answer. The shift on the name. The moment the line turns. In conventional direction you'd give the note at that point in the script; in a prompt field there's nowhere to put it.
The result is that people add words instead. They write "she hesitates, then says" into the script itself, which does two bad things — it burns runtime against a sixty-second cap, and it gets read aloud, because the field is a script rather than a prompt.
An adjective applies to the clip. A trigger applies to a word. That's the whole difference and it's the useful one.
The two modes
Set before you generate, and they're mutually exclusive.
| Auto | Manual | |
|---|---|---|
| What it does | Infers context from the script and renders matching expression and vocal inflection | Takes inline triggers you place at specific words |
| Tagging required | None | A forward slash at any point in the script editor |
| Scope | The whole clip | Wherever you put each trigger |
| Best for | A consistent read where the emotion doesn't shift | Any line where one moment has to land differently from the rest |
Source: Hexcoded Talking Actors documentation, checked September 2026.
Auto is the right default more often than people assume. A thirty-second piece with one consistent register doesn't need tagging, and inferring from the script means the read tracks what you wrote rather than what you tagged.
Manual earns its complexity on anything where the delivery changes within the clip — which for dialogue in short drama is most of it, because a line that doesn't turn isn't doing much.
Why this matters against the cap
Because runtime is the constraint and words are the only thing spending it.
Talking Actors derives duration from script length. There's no slider, no compression setting, and a sixty-second ceiling. Roughly six words reads as three seconds, so every word you write costs runtime you can't get back any other way.
Which makes "she hesitates, then says" a genuinely expensive way to signal a pause. Four words of stage direction is around two seconds of runtime, and it's two seconds of your actor reading the stage direction out loud.
A trigger costs nothing. It's a marker rather than text, so the beat lands without appearing in the word count.
Four words of stage direction costs about two seconds, and gets read aloud. A trigger costs nothing and doesn't.
productWhere to place them
Three positions that do most of the work.
Before the turn. The point where the line changes direction — an answer, a reversal, a concession. That's where a beat reads as thought rather than as a gap.
On the loaded word. A name, a number, an accusation. One word carrying more than the others, marked so the delivery acknowledges it.
At the end, before the cut. Short drama episodes end in suspension rather than resolution, and a beat before the final word is what makes an ending feel unfinished rather than truncated.
Where this falls short. These are placements that make sense from how dialogue works, not positions we've tested against output. What's verifiable is that triggers exist, that they're inline, and that they cost no runtime — where to put them is direction, and direction is yours.
What to do before you generate
Two checks, both free.
Audition the voice. Every actor card with a voice carries a speaker icon, and hearing the read costs nothing. Voice affects perceived pace as much as word count does, and a read that's naturally slower will hit the cap earlier than the estimate suggests.
Run Plan First. It generates a script preview without consuming credits. On a clip written to a specific length, that's where you confirm the estimate before spending anything.
- Two modes. Auto infers from the script across the whole clip; Manual takes inline triggers at specific words
- A trigger is a forward slash placed at any point in the script editor
- An adjective colours the whole generation. A trigger lands on one word. That's the difference
- Duration comes from script length with a sixty-second cap, so every word spends runtime
- "She hesitates, then says" costs about two seconds and gets read aloud. A trigger costs nothing
- Auto is the right default for a consistent read. Manual earns its complexity when the delivery shifts
- Three placements do most of the work: before the turn, on the loaded word, before the cut
- Audition the voice and run Plan First. Both free, both change what you'd write
Two ways. Auto infers context from your script and renders matching expression and vocal inflection across the whole clip, with no tagging. Manual takes inline triggers — a forward slash at any point in the script editor — which place a beat at a specific word.
Scope. An adjective in a prompt applies to the whole generation, so "nervously" colours everything equally. A trigger applies where you put it, which is what you need when one moment has to land differently from the rest of the line.
No. The field is a script rather than a prompt, so anything you write gets read aloud — and it spends runtime against the sixty-second cap. "She hesitates, then says" is about two seconds of your actor reading the direction out loud. Use a trigger instead.
No. A trigger is a marker rather than text, so the beat lands without appearing in the count or the duration estimate. That's what makes it the cheap way to place a pause.
When the emotion doesn't shift within the clip. A thirty-second piece with one consistent register doesn't need tagging, and Auto inferring from the script means the read tracks what you actually wrote rather than what you remembered to tag.
Three positions do most of the work — before the line turns, on the word carrying the weight, and before the final word where an episode ends in suspension. Those are direction rather than rules, so treat them as starting points and audition the voice first, which is free.
Place the beat, don't describe it
Inline emotion triggers land on a specific word without spending runtime, and auditioning a voice costs nothing. Plan First shows you the script before it spends a credit.
Try Talking ActorsMore on prompt structure, per-tool syntax and common mistakes in Prompts.