Text-to-speech voiceover
Generate narration from a script in one prompt — 10 voices, adjustable speed, placed straight on the timeline.
Skip the microphone. Give the agent a script — or just the gist of one — and it generates a natural voiceover and places it on an audio track, ready to mix and caption. Ten voices, adjustable speed, no recording setup.
Prompts to paste
"Add a voiceover: 'Welcome to the tutorial. Here's how to edit a video in one prompt.'"
"Write a 15-second intro voiceover for this product demo and add it at the start — use the nova voice."
"Narrate these three bullet points as a voiceover, slightly faster than normal."
"Redo the voiceover with a deeper voice, same script."
"Generate the voiceover, then balance it against the music."
Notice you don't have to write the script yourself — describe the outcome and the agent drafts it from your project context.
What the agent actually does
The agent sends your text to the text-to-speech engine and adds the resulting audio to a track at the position you choose (default: the start). You control:
- Voice — one of 10:
alloy(default),ash,ballad,coral,echo,fable,nova,onyx,sage,shimmer. - Speed — 0.5x to 2x, default 1. Slower for instruction, faster for social.
- Script length — up to 4,096 characters per generation. For longer scripts, ask the agent to generate it in parts and place them back-to-back.
- Placement —
startSecondsputs it exactly where you want; the clip is a normal audio clip afterwards (trim, split, fade, volume like any other).
Punctuation shapes the delivery — commas add small pauses, periods bigger ones.
When to do it manually
The Text to speech tab in the left panel gives you the same engine with a text box, voice picker, and preview before adding to the timeline — see Text-to-speech panel. Use it when you want to audition voices on a short test line first.
Limits
- TTS requires a configured speech provider on the server. If none is set up, generation degrades gracefully with a clear "not configured" message instead of failing silently.
- 4,096 characters per generation — split longer scripts into parts (the agent can place them sequentially).
- Voices are fixed personalities; there's no custom voice cloning.
- Generated speech is speech like any other: caption it, mix it with auto sound design, cut it on the timeline.
See also
- Text-to-speech panel — the manual path with voice preview
- Auto sound design — mix the voiceover against music
- Auto-caption a video — caption the generated narration