Auto sound design
One prompt balances voice, music, and SFX — sensible dB levels, fades on every clip, and optional muting of camera audio.
A good mix is usually the last thing creators do and the first thing viewers notice. Auto sound design gets you a clean, creator-friendly balance in one prompt: voice up front, music bed underneath, SFX present but not jarring, fades everywhere — without touching a single fader.
Prompts to paste
"Balance the audio — voice clear, music under it."
"Auto sound design, and mute the camera audio since I have a clean voice track."
"Music-forward mix for this cutdown — bring the track up."
"The music is fighting the narration, fix the mix."
"Add fades so nothing starts or ends abruptly."
What the agent actually does
The agent classifies the audio on your timeline as voice, music, or SFX (using asset names and track layout — "bgm", "VO", "sfx", separate tracks vs embedded video audio), then sets a per-clip volume for each type and adds fades:
| Type | Default level |
|---|---|
| Voice / dialogue | 0 dB |
| Music beds | -18 dB |
| SFX | -8 dB |
| Fades | 0.25s in/out on audio clips |
All four numbers are adjustable from your prompt ("music at -12", "longer fades"). Optionally, muteEmbeddedVideoAudio disables the source audio on video clips when a separate voice or uploaded audio track exists — the standard move when you recorded clean audio separately.
It mixes what's already on the timeline — it doesn't fetch music or SFX. To add a track first, ask for one ("add an upbeat music bed") and the agent pulls it from the built-in Freesound library, then mixes it.
When to do it manually
Drag the volume line on any audio clip, or use Properties → Audio for exact dB values and fade handles on clip corners — see Audio editing. Auto sound design plus one or two manual touch-ups is usually faster than mixing from scratch.
Limits
- Classification is heuristic (names, track position) — it doesn't listen to the audio. Name assets clearly ("music-upbeat", "voice-intro") for reliable detection.
- Levels are constant per clip: no LUFS loudness metering and no dynamic ducking that follows the voice moment-to-moment. If music needs to dip only during speech, split the music clip and set levels per section (or ask the agent to).
- It sets clip volumes, not track automation curves.
See also
- Stock sounds and music — where the music and SFX come from
- Text-to-speech voiceover — generate a clean voice track to mix
- Audio editing — manual volume, fades, and audio effects