Auto-caption a video

One prompt adds styled captions timed to speech. Pick from six caption skins, restyle instantly without retranscribing, and fix any line by hand.


Captioning is a one-prompt job in Pluged AI. The agent transcribes your timeline audio and lays styled captions timed to speech — a 5-minute video is captioned in the time it takes to make coffee, and switching styles afterwards is instant because it never re-transcribes.

What you'll make

A fully captioned video in one of six caption skins, with any transcription mistakes fixed by hand, exported with captions burned in.

Prerequisites

  • A video with spoken audio on the timeline (drag it from the Media panel onto the main track)

Step 1: One prompt

Open the Agent panel and type:

"Add captions to this video, TikTok Bold style"

Or let it choose the style for the platform:

"Caption this for a LinkedIn audience"

The agent extracts the audio, transcribes it, and places caption cues on a caption track — each one a real timeline element you can click, drag, and edit.

How the timing works (honest version): transcription is segment-level, not word-stamped. The agent knows when each spoken segment starts and ends, then paces the caption cues inside each segment by syllable weight — longer words get proportionally more screen time. The result reads as captions timed to speech; an individual cue can lead or lag the voice by a fraction of a second. If you need frame-exact timing, import a subtitle file (see Step 5).

Step 2: Pick a skin

Six caption skins ship with the app. Each sets the typography, background box, and screen position — nothing else (no strokes, no glow, no animations):

SkinLookBest for
TikTok BoldLarge bold white text on a dark rounded backingTikTok, Reels, Shorts
Minimal CleanPlain white text, no background boxInterviews, tutorials, clean brands
Boxed ContrastBold text in a near-solid black boxBusy or bright footage, accessibility
Editorial HighlightSerif type on a left-aligned dark railDocumentary, premium editorial
Neon PopBright text on a deep blue boxGaming, hype edits, music
Documentary LowerSmall quiet serif low in the frameLong-form docs, staying out of faces

Name the skin in your prompt, or describe the vibe ("something quiet that stays out of my face") and the agent picks.

Step 3: Restyle without retranscribing

Changing your mind costs nothing:

"Switch the captions to Boxed Contrast"

This restyles the existing caption elements in place — the transcript is cached, so no re-transcription, no waiting, and any text corrections you've made are preserved. Try two or three skins against your footage; it's the cheapest taste decision in the app.

Step 4: Fix mistakes by hand

Transcription is good, not perfect — names, jargon, and mumbled lines miss. Captions are ordinary timeline elements:

  • Wrong word: double-click the caption in the preview or timeline, retype it, press Enter
  • Timing off: click the caption clip and drag it left/right; drag its edges to lengthen or shorten
  • Cue splits awkwardly: ask the agent — "merge the captions at 0:42" or "split that long caption into two"
  • Custom look: select a caption and use the Properties panel to override font, size, color, or position for that cue

Step 5: Already have subtitles?

If you have an .srt or .ass file (from a human transcriber or another tool), import it from the Captions tab instead of generating — you get its exact timing and text, and can still restyle with any skin. Full details in the captions deep-dive.

Step 6: Export

ExportMP4 (H.264)High. Captions are burned into the rendered video — what you see in the preview is what viewers get.

Make it yours

"Caption this in Spanish" — the agent passes a language to transcription; leave it out and the language is auto-detected

"Caption it, then delete the caption cues over the intro — I want the first 5 seconds clean"

"Caption it, then move the captions higher — they clash with my lower-third"

The cached transcript also unlocks bigger asks in the same session: "now cut a 30-second highlight of the strongest argument" reuses the transcript to pick moments. See transcript highlights.

Troubleshooting

  • Captions drift from the voice in one section. Segment-level timing paces cues within each spoken segment, so a fast or mumbled stretch can drift. Drag those caption clips into place, or import an SRT for exact timing.
  • No captions appeared. The clip needs audible speech — check the clip isn't muted (Properties → Audio) and the volume isn't at zero, then ask again.
  • Wrong language detected. Say "re-caption this — the audio is French." Re-generating replaces the caption track.
  • Style change didn't apply. Skins restyle existing captions — if there are none yet, the agent needs to caption first. "Caption this with Neon Pop" does both in one go.
  • Captions collide with other text. Ask "move the captions clear of the title" — the agent's own quality checks flag overlapping text, but manual titles added afterwards can reintroduce it.

See also

Community