Extract best moments from video
Turn long footage into a tight highlight reel — AI-ranked best ranges when available, audio energy as fallback.
Turn 20 minutes of raw footage into a tight highlight reel with one prompt. The agent keeps the strongest moments, deletes the rest, and closes the gaps — a job that takes an hour of scrubbing by hand takes about a minute here.
Copy-paste prompts
"Extract a 60-second highlight reel."
"Watch the footage first, then cut it down to the best 30 seconds."
"Make a 45-second highlight reel with longer moments — around 8 seconds each."
"Turn this into a 30s vertical highlight with captions."
The last prompt is the real workflow: one delegation triggers analysis, extraction, reframing, captions, and the agent's own quality checks in a single run.
What the agent actually does
Here's what happens when you ask for the best moments:
- Prefers Pegasus-ranked ranges. If visual analysis has run, the tool keeps the AI-ranked
bestRanges— moments chosen by a model that actually watched the footage. - Falls back to audio energy. Without cached analysis, it keeps the highest-energy windows on the main track (louder, more animated speech scores higher).
- Compacts the timeline. Selected moments stay, everything else is deleted, and main track, overlays, and audio ripple together to keep sync.
Every change lands on the real timeline and is undoable.
Parameters
| Parameter | Default | What it controls |
|---|---|---|
targetDuration | 30 seconds | Total length of the highlight reel |
windowSeconds | 4 seconds | Approximate length of each retained moment |
There is no "number of moments" parameter — ask for a duration instead ("the best 90 seconds"), and shape pacing with the window ("longer moments, around 8 seconds each" for breathing room; shorter for rapid-fire montages).
Get better selections
Run visual analysis first — or just ask for both in one prompt:
"Analyze the footage, then extract the best 60 seconds."
With analysis cached, selection is based on what's on screen (reactions, action, key demonstrations), not just how loud it was.
For speech-heavy content — podcasts, interviews, talks — use transcript highlights instead. It selects by what was said rather than how it sounded:
| Tool | Best for | Selects by |
|---|---|---|
| Extract best moments | Action, events, vlogs, gameplay | Pegasus-ranked visuals, else audio energy |
| Extract transcript highlights | Podcasts, interviews, presentations | High-signal transcript segments |
Iterate conversationally
The first cut is a draft. Redirect it in plain language:
"You cut my favorite part around 2:30 — bring that back."
"Too choppy. Redo it with fewer, longer moments."
"Make it 90 seconds instead."
Or take over manually: drag clip edges on the timeline to fine-tune in/out points.
When to do it manually
Scrub the footage, split at the moments you want (S at the playhead), delete the rest, and close gaps — see Timeline editing. Sensible for short clips where you already know the exact moments.
Limits
- Result duration is close to the target, not exact — clean cut points win over hitting the number precisely.
- Audio-energy fallback favors loud moments; quiet-but-great content needs visual analysis or transcript highlights.
- The tool works on the main track; run it before adding B-roll and overlay cards, not after.
- Visual analysis is metered on the free plan — see Plans and limits.
See also
- Extract transcript highlights — selection by spoken content
- Visual analysis — where
bestRangescome from - One prompt, full edit — the full delegation workflow
- Podcast to clip — highlight workflow for talking-head footage