How the AI agent works
Hand the agent one outcome — it plans, cuts, captions, and reviews its own edit on your real timeline. Every step is undoable.
The Pluged AI agent is built for delegation, not micromanagement. You give it one outcome-level prompt — "turn this into a 30-second TikTok with a hook and captions" — and it plans the edit, makes every cut on your real timeline, then watches its own work and fixes what's wrong before it tells you it's done. What used to be an afternoon of trimming, captioning, and second-guessing becomes one prompt and a few conversational refinements.
You stay in charge the whole time: every change is a normal timeline command you can undo with Cmd/Ctrl+Z, and you can take over manually at any point.
Start with one prompt
Describe the video you want, not the steps to get there:
Turn this into a 30-second TikTok with a strong hook and captions.
Cut this interview down to the three best moments, make it vertical, and add captions.
Make this screen recording into a 60-second product demo that ends with a call to action.
Edit this talking-head clip into something engaging for Reels — you pick the moments.
Then iterate conversationally:
Tighten the middle — it drags after the hook.
Use a bolder caption style and make the whole thing 5 seconds shorter.
What happens on your first message
The agent routes by what's already in your project instead of opening with questions:
- Footage on the timeline + a full-edit request — it states its thesis and target length in one line, then starts working.
- Imported but unplaced footage — it analyzes the footage first, proposes one route in a single line, and starts.
- An empty project — it asks whether to wait for your import or build from stock footage.
- A question ("what's on the timeline?") — it answers from the project state, no tools.
It asks at most two questions per request, and only when the answer genuinely changes the edit (for example, a 3-minute YouTube explainer vs. a 30-second reel). Everything else it defaults sensibly — aspect ratio from your canvas, 30–60 seconds for reels, captions on for talking heads — and tells you the default it picked.
The storyboard: the plan you can see
Before its first cut on any full edit, the agent writes a storyboard — a one-line thesis, a target length, and 3–6 beats (hook, context, payoff, CTA) with the source footage ranges each beat comes from. It appears as a collapsible card above the chat, beats flip from planned to done as the agent works, and it persists across turns and reloads. If you change direction mid-edit, the storyboard updates first, then the timeline follows.
See The storyboard card for how to read and steer it.
What the agent actually does
For a full edit, the agent follows a fixed recipe rather than improvising:
- Understand — analyzes your footage and transcript so it's directing footage it has actually looked at.
- Plan — writes the storyboard: thesis, target length, beats with source ranges.
- Select and order — keeps whole-thought segments that serve the thesis (a 30-second reel is 3–6 clips, not 15 fragments), opening with the strongest line rather than chronological order.
- Format and pace — sets the right canvas, reframes for faces or subject, and adds motivated motion: zoom punches at cut points, B-roll on the concrete thing being spoken, an overlay card when a stat or name is said.
- Captions, audio, style — captions timed to speech in one consistent skin, a music bed ducked under speech that runs the full length of the cut, and exactly one template or style pack.
- Watch and fix — reviews its own edit before finishing (see below).
Each tool call appears as a tool row in the chat: the action in plain English, a status (running, applied, or error), and expandable details of what changed. Every tool call lands as a regular timeline command, so a single Cmd/Ctrl+Z undoes any step — the agent never has a separate, un-undoable edit path. A handful of tools are read-only (analysis and quality checks) and never touch the timeline.
How it checks its own work
Before finishing a full edit, the agent runs a mandatory sequence: a creator-basics critique (plus a deterministic lint for overlapping text, text out of frame, music-bed coverage, and machine-gun cuts), then a check of the hard constraints (duration, captions, canvas), then an AI watch-through — the agent renders your cut and has an AI watch it like an editor, returning problems, fixes, and a ship/no-ship verdict.
A quality gate enforces this: if the agent tries to conclude while its latest check is failing, the loop bounces it back to fix the issues (up to twice per run). It can't just describe the edit as done — the checks have to pass.
Full details in How the agent checks its work.
Point the agent at exactly what you mean
Type @ in the chat composer to anchor your prompt to something specific:
@project— the whole project@selection/@clip— the clips you currently have selected@range(0:05-0:12)— a specific time range- Or pick any timeline clip from the
@menu
Trim @selection to just the punchline and add a zoom punch on the cut.
Make @range(0:05-0:12) the opening hook.
Staying in control
By default the agent auto-applies edits — each change lands immediately and is undoable. Two ways to slow it down:
- Review before applying — a persistent toggle in the agent settings menu (the small status text under the composer reads "Auto-applies edits" or "Review before applying"). With it on, the agent's changes wait for your approval instead of applying live.
- Ask for a plan — say "propose a plan first" and the agent drafts a plan card you can read and run, or discard. This is a one-off; the toggle is the standing preference.
And at any moment you can just stop it, undo, and edit manually — the agent works on the same timeline you do.
Choosing a model
The model picker in the agent panel offers Claude Sonnet (the default — fast and strong for almost every edit) and Claude Opus (heavier reasoning for the hardest edits). On the free plan the picker also shows an AI usage bar — the percentage of your monthly budget used, never raw tokens — with an upgrade prompt at 100%. See Plans and limits.
Recipes
When the agent nails an edit, you can save its tool sequence — the exact calls and parameters, from a draft or a plan — as a named recipe and replay it on new footage in any project. See Saving and replaying recipes.
What the agent can do
In one honest summary: the agent can analyze footage and transcripts; assemble edits from best moments, transcript highlights, or your timestamp picks; cut silences and filler words; reframe for vertical; apply creator templates, style packs, and caption skins; generate captions, hooks, text, graphics, and animated overlay cards; fetch stock footage, music, and sound effects; generate TTS voiceover; mix sound; add transitions, effects, and zoom punches; manage scenes and canvas; and check its own work. That's 41 distinct capabilities in total — the complete list is in the Agent capabilities reference.
It does not see your file system or run code, and it can't import files from local paths — it works with media already in your project plus what its stock, sound, and voiceover tools can fetch.
When to do it manually
Everything the agent does maps to the manual editor: trim and arrange on the timeline, style text in the properties panel, and add captions from the Captions tab (captions deep dive). The agent is fastest for the first 80%; manual editing is often fastest for the last pixel-level 20%.
Limits
- Transcription is segment-level, not word-timestamped — captions are timed to speech with syllable-weighted pacing, not per-word karaoke.
- Reframing is heuristic, based on cached visual analysis — not pixel-level face tracking.
- The AI watch-through skips timelines longer than 120 seconds (the other checks still run).
- Stock media, sounds, TTS, and video understanding depend on configured provider keys; when one isn't configured, the tool reports it and the agent works with what it has.
- Free-plan usage is metered monthly; at 100% new agent requests are blocked until the next cycle (in-flight runs finish).