Extract transcript highlights

Cut podcasts and interviews down to their strongest spoken moments — transcript-scored, goal-aware, ripple-safe.


Cut an hour-long conversation down to the parts worth clipping. Transcript highlights scores every spoken segment for signal, keeps the winners up to your target duration, and removes the rest with sync intact — the fastest path from podcast or interview to a postable clip.

Copy-paste prompts

"Extract the best 30 seconds based on the transcript."

"Keep the segments with the strongest arguments about pricing."

"Caption this first, then pull the funniest 60 seconds."

"Extract highlights about the product launch, then reframe it vertical with captions."

What the agent actually does

Here's what happens when you ask for transcript highlights:

  1. Reads the cached scene transcript (created when the agent captions your video).
  2. Scores each segment: information density, signal words ("because", "reason", "result"), meaningful punctuation, grammatical completeness, and match against your stated goal.
  3. Keeps the top-scoring segments in original order, up to the target duration, with a little padding around each so speech isn't clipped mid-breath.
  4. Deletes everything else and ripples main track, overlays, and audio together — captions stay in sync with the cut.

Parameters

ParameterDefaultWhat it controls
targetDuration30 secondsTotal retained speech
goalnoneTopic or outcome to bias selection ("strongest arguments", "funniest moments")
paddingSeconds0.18 secondsBreathing room around each kept segment

Prerequisite: a transcript

This needs the cached scene transcript that captioning creates. If you haven't captioned yet, just ask for both — the agent runs them in order:

"Generate captions, then extract highlights about marketing."

Where it fits

Combine withResult
Auto captionsCreates the transcript this tool scores
Filler removal"Remove fillers, then extract highlights" — cleaner input, cleaner cut
Smart reframeVertical crop after the cut
Caption skinsRestyle captions on the finished clip

When to do it manually

Read the transcript in the Captions panel, note the timestamps you want, and split/delete on the timeline. Worth it only when you need editorial judgment the scoring can't capture (a specific narrative arc, a legal-safe cut).

Limits

  • Segment-level, not word-level. Transcription is segment-based (segments are typically 2–10 seconds), so cuts land on segment boundaries — you can't surgically keep half a sentence.
  • Keyword-goal matching. The goal biases toward segments containing related words; it's not full semantic search.
  • Clear speech works best. Crosstalk and heavy background noise score unpredictably.
  • Result duration is approximate — clean segment boundaries win over hitting the target exactly.

See also

Community