Auto-caption a video
One prompt adds styled captions timed to speech. Pick from six caption skins, restyle instantly without retranscribing, and fix any line by hand.
Captioning is a one-prompt job in Pluged AI. The agent transcribes your timeline audio and lays styled captions timed to speech — a 5-minute video is captioned in the time it takes to make coffee, and switching styles afterwards is instant because it never re-transcribes.
What you'll make
A fully captioned video in one of six caption skins, with any transcription mistakes fixed by hand, exported with captions burned in.
Prerequisites
- A video with spoken audio on the timeline (drag it from the Media panel onto the main track)
Step 1: One prompt
Open the Agent panel and type:
"Add captions to this video, TikTok Bold style"
Or let it choose the style for the platform:
"Caption this for a LinkedIn audience"
The agent extracts the audio, transcribes it, and places caption cues on a caption track — each one a real timeline element you can click, drag, and edit.
How the timing works (honest version): transcription is segment-level, not word-stamped. The agent knows when each spoken segment starts and ends, then paces the caption cues inside each segment by syllable weight — longer words get proportionally more screen time. The result reads as captions timed to speech; an individual cue can lead or lag the voice by a fraction of a second. If you need frame-exact timing, import a subtitle file (see Step 5).
Step 2: Pick a skin
Six caption skins ship with the app. Each sets the typography, background box, and screen position — nothing else (no strokes, no glow, no animations):
| Skin | Look | Best for |
|---|---|---|
| TikTok Bold | Large bold white text on a dark rounded backing | TikTok, Reels, Shorts |
| Minimal Clean | Plain white text, no background box | Interviews, tutorials, clean brands |
| Boxed Contrast | Bold text in a near-solid black box | Busy or bright footage, accessibility |
| Editorial Highlight | Serif type on a left-aligned dark rail | Documentary, premium editorial |
| Neon Pop | Bright text on a deep blue box | Gaming, hype edits, music |
| Documentary Lower | Small quiet serif low in the frame | Long-form docs, staying out of faces |
Name the skin in your prompt, or describe the vibe ("something quiet that stays out of my face") and the agent picks.
Step 3: Restyle without retranscribing
Changing your mind costs nothing:
"Switch the captions to Boxed Contrast"
This restyles the existing caption elements in place — the transcript is cached, so no re-transcription, no waiting, and any text corrections you've made are preserved. Try two or three skins against your footage; it's the cheapest taste decision in the app.
Step 4: Fix mistakes by hand
Transcription is good, not perfect — names, jargon, and mumbled lines miss. Captions are ordinary timeline elements:
- Wrong word: double-click the caption in the preview or timeline, retype it, press Enter
- Timing off: click the caption clip and drag it left/right; drag its edges to lengthen or shorten
- Cue splits awkwardly: ask the agent — "merge the captions at 0:42" or "split that long caption into two"
- Custom look: select a caption and use the Properties panel to override font, size, color, or position for that cue
Step 5: Already have subtitles?
If you have an .srt or .ass file (from a human transcriber or another tool), import it from the Captions tab instead of generating — you get its exact timing and text, and can still restyle with any skin. Full details in the captions deep-dive.
Step 6: Export
Export → MP4 (H.264) → High. Captions are burned into the rendered video — what you see in the preview is what viewers get.
Make it yours
"Caption this in Spanish" — the agent passes a language to transcription; leave it out and the language is auto-detected
"Caption it, then delete the caption cues over the intro — I want the first 5 seconds clean"
"Caption it, then move the captions higher — they clash with my lower-third"
The cached transcript also unlocks bigger asks in the same session: "now cut a 30-second highlight of the strongest argument" reuses the transcript to pick moments. See transcript highlights.
Troubleshooting
- Captions drift from the voice in one section. Segment-level timing paces cues within each spoken segment, so a fast or mumbled stretch can drift. Drag those caption clips into place, or import an SRT for exact timing.
- No captions appeared. The clip needs audible speech — check the clip isn't muted (Properties → Audio) and the volume isn't at zero, then ask again.
- Wrong language detected. Say "re-caption this — the audio is French." Re-generating replaces the caption track.
- Style change didn't apply. Skins restyle existing captions — if there are none yet, the agent needs to caption first. "Caption this with Neon Pop" does both in one go.
- Captions collide with other text. Ask "move the captions clear of the title" — the agent's own quality checks flag overlapping text, but manual titles added afterwards can reintroduce it.
See also
- Caption skins — every skin in detail
- Captions deep-dive — SRT/ASS import, manual generation, track behavior
- Make a TikTok from raw footage — captions as part of a full edit