Text-to-speech voiceover

Generate narration from a script in one prompt — 10 voices, adjustable speed, placed straight on the timeline.


Skip the microphone. Give the agent a script — or just the gist of one — and it generates a natural voiceover and places it on an audio track, ready to mix and caption. Ten voices, adjustable speed, no recording setup.

Prompts to paste

"Add a voiceover: 'Welcome to the tutorial. Here's how to edit a video in one prompt.'"

"Write a 15-second intro voiceover for this product demo and add it at the start — use the nova voice."

"Narrate these three bullet points as a voiceover, slightly faster than normal."

"Redo the voiceover with a deeper voice, same script."

"Generate the voiceover, then balance it against the music."

Notice you don't have to write the script yourself — describe the outcome and the agent drafts it from your project context.

What the agent actually does

The agent sends your text to the text-to-speech engine and adds the resulting audio to a track at the position you choose (default: the start). You control:

  • Voice — one of 10: alloy (default), ash, ballad, coral, echo, fable, nova, onyx, sage, shimmer.
  • Speed — 0.5x to 2x, default 1. Slower for instruction, faster for social.
  • Script length — up to 4,096 characters per generation. For longer scripts, ask the agent to generate it in parts and place them back-to-back.
  • PlacementstartSeconds puts it exactly where you want; the clip is a normal audio clip afterwards (trim, split, fade, volume like any other).

Punctuation shapes the delivery — commas add small pauses, periods bigger ones.

When to do it manually

The Text to speech tab in the left panel gives you the same engine with a text box, voice picker, and preview before adding to the timeline — see Text-to-speech panel. Use it when you want to audition voices on a short test line first.

Limits

  • TTS requires a configured speech provider on the server. If none is set up, generation degrades gracefully with a clear "not configured" message instead of failing silently.
  • 4,096 characters per generation — split longer scripts into parts (the agent can place them sequentially).
  • Voices are fixed personalities; there's no custom voice cloning.
  • Generated speech is speech like any other: caption it, mix it with auto sound design, cut it on the timeline.

See also

Community