Cindy Zhu.
← all free guides
Content creation ChatGPT

Edit your videos with AI: the good-take workflow

Hey, it's Cindy ๐ŸŒฑ This is the full setup for editing with ChatCut, the plugin that turns your desktop agent into a real editing timeline. It genuinely takes the mechanical work off your plate. It does not take the taste off your plate, and this guide is honest about which is which.
โš ๏ธ Read this before you install anything. This is a developer-style install: it runs in the desktop app, not ChatGPT on the web or your phone, and it needs Git available on your machine. If that sentence worries you, use ChatCut's own web editor instead. You get the same editor and the Skills feature, just without driving it from a chat. Worth knowing either way: both agent pieces are still pre-1.0, the desktop app has no auto-update, and the two open bugs on their public repo are both on the Windows path.


what you're actually getting

What ChatCut is, and what it can do ๐ŸŽฌ

It's a cloud video editor with a real multi-track timeline, and the plugin lets an AI agent drive it for you. The editing doesn't happen inside the chat: your agent sends instructions to ChatCut, and the timeline lives in your browser at app.chatcut.io. That matters because you can always open the project and finish by hand.

The same plugin works in the ChatGPT desktop app, Claude Code, Cursor, Grok and Kimi, so you're not locked to one assistant.

โœ‚๏ธ Cut by editing text
Delete words in the transcript and the matching video is removed and the gap closed. This is the Descript-style workflow, and it's the thing it does best.
๐Ÿ’ฌ Captions
Styled presets, per-word styling, and karaoke-style highlighting of the word being spoken. Translation into other languages too.
๐ŸŽž๏ธ B-roll and multicam
Full-screen cutaways or rounded picture-in-picture, from your own library or stock. Multicam syncs angles by their audio.
โœจ Motion graphics as code
Unusually, it writes the graphic as code, so every piece stays editable afterwards instead of being a flat video.
๐Ÿ”Š Voice, music and cleanup
Generated voiceover, music and sound effects, plus background noise removal that leaves the picture untouched.
๐Ÿ“ค Export
MP4 or WebM to 1080p in the browser, 4K through their desktop app, SRT subtitles, and an XML timeline for Premiere or DaVinci Resolve.
โš ๏ธ Know the export limits before you build a workflow on it. The timeline export is one format only (the older FCP7 XML), and it drops captions, effects, transitions and motion graphics on the way out, so treat it as a rough handoff, not a faithful one. Subtitles come out as SRT or plain text, never VTT. Audio-only exports are MP3. The CapCut draft handoff is real but only works from their desktop app, not through the plugin.
๐Ÿ’ก Careful what you install. There are unrelated open-source projects called OpenChatCut that come up in searches. ChatCut itself is a paid cloud product, not an open-source app, and the plugin is the only open part.
step one

Install the plugin ๐Ÿ”Œ

Three routes, all documented by ChatCut. Pick one:

  1. โšก The one-liner (fastest) Open the ChatGPT desktop app and paste this into any task. It installs itself in about a minute.
    paste into the desktop app
    /goal Read chatcut.io/chatgpt to install the ChatCut plugin and set up a new task for me.
  2. ๐Ÿ”Ž From the plugin store Open the plugin store in the desktop app, search ChatCut, click Install.
  3. ๐Ÿ›  Manually Plugins in the sidebar โ†’ Create โ†’ Add marketplace โ†’ paste https://github.com/ChatCut-Inc/agent-plugin.git โ†’ open the Personal tab โ†’ find ChatCut โ†’ Install.
  4. ๐Ÿ”‘ Sign in when it asks It opens ChatCut's authorisation page in your browser. Sign in and allow. Never paste passwords or tokens into the chat itself.
๐Ÿšจ The one that catches everyone: the chat you install in can never edit. ChatCut's docs are blunt about it, the install conversation cannot reach the newly installed tools even after login succeeds. Start a brand-new conversation, then type @chatcut to wake it. If your first edit request does nothing, this is why.
๐Ÿ’ฅ Do not drag a video onto the embedded editor panel. ChatCut says it "can crash the desktop app," and it's a host-side bug they've reported to OpenAI. Add media through the chat column, or inside ChatCut via My Assets โ†’ Upload. Uploads through the plugin get normalised to 1920px and 30fps, four files at a time, so don't bulk-dump a folder.

before you start

What it costs ๐Ÿ’ณ

๐Ÿ–ฅ The desktop app Mac or Windows. ChatCut says every plan including Free works. Realistically, a long edit is a long agent session, so a paid tier is the comfortable floor.
๐ŸŽฌ A ChatCut account Their docs describe a free tier with the core editor and a one-time credit balance; their pricing page lists no free tier and one of their comparison pages says there is none. If you get it, it carries a 60-minute cumulative export quota that never resets. Not per month. Sixty, ever. The good news: they state plainly that there is no watermark on exports, on free or any plan.
๐Ÿ’ฐ Credits buy generation, not editing $25 to $160 a month for 100 to 800 credits. But cutting, transcription, captions and exporting cost zero credits. Credits go on generated video, music and voice.
๐Ÿ“ค Export, and a one-way door MP4 or WebM up to 1080p on the web, 4K only via their desktop app (local render, not a paid upgrade). You can send XML to Premiere or Resolve, but you can never import an existing edit back in.
๐Ÿ’ธ Check the credit maths yourself, because the headline numbers are 480p. Their pricing page offers 100 credits as roughly 263 seconds of generated video. That rate is the 480p price. At 1080p their own billing table works out closer to 45 seconds for the same 100 credits. The free starting balance of 20 credits is about 9 seconds of 1080p generated video, once, ever, and it does not renew. None of that touches ordinary editing, which is free, so this only bites if you lean on generated B-roll. Also: talking to the agent costs credits too, and their docs say longer conversations cost more, so a tight brief is cheaper than a long negotiation.

do this while filming

Mark your good takes out loud ๐ŸŽฌ

You will do the same line five times. The agent can see all five and has no idea which one you liked. So tell it while you're still rolling.

After a take you're happy with, say a marker out loud in your normal voice, then carry on. ChatCut's own blog describes beta users doing exactly this, saying "ChatCut, mark this as director's favorite," then adding to the prompt: "Never use the parts addressing 'ChatCut' in the final video, but use them as logging notes." That second half matters, so include it.

๐ŸŽž๏ธ This isn't a hack, it's an old film habit. It's the circle take: the script supervisor circles the director's preferred take on the script pages so the edit suite knows what to use. You're doing the same thing, just leaving the note in the audio instead of on paper.
๐ŸŒฑ If you already write scripts, there's a second route. ChatCut aligns your footage to a written script and surfaces the strongest version of each line across every take. So you can paste your script in and pick from the attempts it shows you. Use the spoken marker, the script, or both. The marker wins when you improvise; the script wins when you don't.

the order matters

Cut the speech first, everything else second โœ‚๏ธ

This is the single most useful rule I found, and it's in ChatCut's own guidance: the speech timing set by your A-roll cut anchors everything downstream. If you ask for cuts, captions, graphics and music in one prompt, you cannot tell which decision went wrong when it comes back odd.

So go in passes. Their own copy-paste prompts, in order:

1 ยท the first cut
Use ChatCut to make a clean first cut of this talking-head video. Remove repeated takes, filler words, obvious restarts, and long empty pauses, but preserve the speaker's meaning and natural cadence. Keep the strongest opening line, add readable captions, and leave all media and edits on an editable timeline. Do not add generated footage or music yet.
2 ยท tighten it
Remove repeated takes, filler words, obvious restarts, and unnecessary pauses. Keep the meaning intact, preserve any pause that helps the point land, and make the delivery feel natural.
3 ยท captions
Add word-level captions and highlight the word currently being spoken. Keep the text readable on a phone, away from the speaker's face, and consistent throughout the video.
4 ยท sound and movement, last
Add subtle sound effects, transitions, and punch-in zooms only where they connect ideas or emphasize an important line. Keep the existing cut unchanged and avoid an effect on every sentence.

One more worth keeping, for turning a long recording into shorts:

5 ยท long-form into shorts
Find three self-contained highlights in this recording. Turn each into a separate 9:16 short. Open with the strongest complete line, preserve enough context for each idea to stand alone, keep each timeline editable, and add phone-readable captions. Do not invent quotes or cut a sentence in a way that changes its meaning.

making it look like you

Teach it your look ๐ŸŽจ

Here's where I have to correct something I believed myself: you cannot hand it links to your old videos and have it learn your style. That isn't a feature. What exists is better documented and easier:

  1. ๐Ÿ–ผ Build a Design Style Open Design Style โ†’ Create Design Style. It takes colours, fonts, logos, reference images and written rules. Screenshot the frames from your own videos you want it to copy and upload those. No scraping tools, no ffmpeg, just screenshots.
  2. ๐Ÿ”ค Pick fonts from its catalogue Not from your machine. ChatCut warns that system fonts like Arial or Helvetica fall back unpredictably between computers. Custom fonts attach to your Design Style.
  3. ๐Ÿ”Š Use the built-in sound library first There's a sound effects library, and their own guidance is to use it before generating anything, because generating costs credits and the library doesn't.
  4. ๐Ÿ’พ Save it as a Skill Once a workflow works, open a project, click the book icon under the message field, and choose Save this editing process as a Skill. It comes back under My Skills in your other projects.
โš ๏ธ Design Styles mainly shape motion graphics, not your cutting rhythm. And Skills are documented inside ChatCut's own editor, not the chat plugin. So if the reusable-workflow part is what you're here for, drive it from ChatCut's editor rather than from a chat window.

the honest part

What it does well, and what still needs you ๐Ÿ‘ฉโ€๐Ÿ’ป

โœ… Genuinely good The mechanical cut. Pulling repeated takes, closing dead air, trimming false starts, captioning, placing effects where you said. In the one detailed public test, 51 minutes of raw footage became a 22-minute first cut, with about 25 to 30 minutes of agent runtime. This is the hours you were losing.
๐Ÿง  Still yours Which take had the energy. Whether a section should go entirely. Whether a pause is dead air or a beat. ChatCut's own blog says it plainly: "Speed gets you to a first draft, human judgment determines what actually goes live."

Where it actually breaks. Independent coverage is genuinely thin, so this is the one substantive outside critique that exists, and it matches my own experience: the agent "can still struggle to interpret more subjective stylistic choices, pacing nuances, or complex narrative intents," and auto-picked B-roll "can feel repetitive and generic." ChatCut itself concedes that transcription "misreads proper nouns, brand names, technical terms, and fast or accented speech," which matters if you say tool names on camera.

The failure worth bracing for, from an editor who hit it: "I asked an editing agent to 'make this punchier,' got a tight cut back, and then realized it had removed the one sentence that made the whole point make sense."

๐Ÿ“Š About the "80%" figure, mine included. It's fair for the mechanical work on a solo talking head, which is the easiest case there is. It's not 80% of the edit. Adobe surveyed 16,000+ creators this year: 57% said AI outputs need moderate or extensive editing before they're shareable, and 85% said the final creative call should always stay with the creator. Plan on directing it, not delegating to it.
๐Ÿ” Expect several rounds before it's predictable. These agents are non-deterministic, so the same prompt can behave differently twice. That's not you doing it wrong. Correct it in plain words, run again, and keep your correction notes, because by round three you've effectively written your style guide. Just remember the meter is running while you do it.


the fun part

Motion graphics and little animations ๐ŸŽจ

This is the part most people miss. ChatCut writes motion graphics as code, which means the animated stat, title card or lower third it makes for you stays editable afterwards. You can ask for a change instead of starting again. Ask in one pass, after the cut is locked, so the graphic lands on timing that won't move.

โœจ an animated stat card
Add a motion graphic to the timeline at [the moment I say the number], on top of the existing footage, and do not change the cut underneath it.

It's a stat card that counts up to [87%] with the label [of viewers never make it past 3 seconds]. Vertical 1080 by 1920, about [3] seconds, positioned in the [upper third] so it never covers my face.

Style: one accent colour [#D97757] on a [cream] card, everything else neutral, one clean sans-serif. The number counts up with an ease-out, the label fades in slightly after it, and the whole card leaves the way it came in. No bouncing, no spinning, nothing pulsing while it sits there.

Show me a preview frame before you render it, and keep the graphic editable so I can change the number later.
๐Ÿ’ก Why "nothing pulsing" is in there. The quickest way to make an overlay look amateur is idle movement: things that wobble or breathe while they wait. Good motion moves with purpose, then stops. Ask for one accent colour too, because a graphic with five colours reads as decoration instead of information.

Your own little mascot animation ๐Ÿฆ€

A small looping character is a cheap way to make a video feel like yours, and you don't need a design tool. Ask your agent to draw it in code as pixel art, so every version afterwards is a number change rather than a re-roll. You need Python and one free library (pip3 install Pillow), and your agent can set that up for you.

๐Ÿฆ€ a looping pixel mascot
Make me a small animated pixel mascot in Python using Pillow. Draw every frame in code from square blocks. Do not generate images with AI.

Character: [a round orange cat with two dot eyes]. Action: [typing on a laptop, looking pleased with itself].

Rules: build it from chunky pixel blocks so it reads as 8-bit, and write one function that draws the character with arguments for the parts that move, then call it once per frame so the character never changes size or colour between frames. Every frame is transparent, with no background. Pick one ground line and keep the feet on it, so it never floats. Around 16 to 24 frames for a simple loop. Export one looping GIF with a transparent background, saved as [name].gif.

Then show me the first frame and tell me the frame count. I'll tell you what to change and you adjust the numbers and export again.

Expect the first working version in about ten minutes, then a couple of minutes per tweak. Drop the GIF over your footage in ChatCut, or in whatever you edit in.


free and worth knowing

The free tools that do the boring parts ๐Ÿ› ๏ธ

You don't need any of these to use ChatCut, but they're free, they run on your own machine, and an AI agent can drive all of them. This is the layer that costs you nothing and saves hours.

  1. โœ‚๏ธ auto-editor, the silence cut Cuts dead air and filler words from a recording, and can export a timeline straight to Premiere, Resolve or Final Cut instead of a finished video. Install with brew install auto-editor, then run auto-editor video.mp4 --margin 0.2s. Public domain, genuinely free.
  2. ๐ŸŽง ffmpeg, for everything else Converting, cropping to 9:16, normalising loudness, burning in captions, exporting a transparent overlay. Nobody remembers the syntax, which is exactly why handing it to an agent works so well. Install with brew install ffmpeg-full.
  3. ๐Ÿ“ Transcription on your own machine The full ffmpeg build can now transcribe audio to an SRT file by itself, and brew install whisper-cpp gives you a faster one built for Apple Silicon. Free, offline, and nothing gets uploaded.
  4. ๐Ÿช„ rembg, for cutouts Removes the background from an image in one command: pip install "rembg[cli]", then rembg i input.png output.png. Handy for thumbnails and overlays.
โš ๏ธ The trap in every captions tutorial. Plain brew install ffmpeg ships without the caption filters, so the usual burn-in command fails with a confusing "no such filter" error. Install brew install ffmpeg-full instead. It doesn't replace your existing ffmpeg automatically, so either tell your agent to use the full path it prints after installing, or add it to your PATH with export PATH="/opt/homebrew/opt/ffmpeg-full/bin:$PATH". To check which one you have, run ffmpeg -filters | grep subtitles: if nothing comes back, captions will fail.
๐Ÿ’ก Ask before you install. Paste this into your agent: "I want to cut the silence out of [file], burn in captions from an SRT, and export a 9:16 version. Tell me which of ffmpeg-full and auto-editor I need, install what's missing, then do it one step at a time and show me the result after each step."
when you outgrow it

The ceiling, and what's past it ๐Ÿ“ˆ

ChatCut's motion graphics are genuinely good for captions, simple animations and basic charts, and you can edit the generated pieces afterwards. What it can't do is bind a chart to a real dataset or give you frame-exact control.

If you need that, the step up is building the graphic as code, where a chart reads from a real dataset and you control every frame. Two options worth knowing, and the licensing differs sharply. Remotion is not open source, it's source-available under a proprietary licence, but it's genuinely free for an individual, commercially, forever. That only changes at four or more people. HyperFrames is Apache-2.0 with no headcount threshold at all.

The learning curve is smaller than it was. Both now have coding-agent paths, so you can prompt your way to a first render rather than writing React from scratch. What you get back is still code, so every refinement means reading or re-prompting it. Reach for this only for the recurring branded asset, the chart or title treatment you'll rebuild every month. It's an overlay on your edit, not a replacement for it.



the question everyone asks

Does it replace Descript? ๐Ÿค”

For the job most creators use Descript for, cutting a talking head down from the transcript, removing filler and dead air, and getting a tight first cut, yes, it genuinely does. And it adds things Descript doesn't have: an outside AI agent that can run the whole cut for you, motion graphics that stay editable as code, and generation of video, images and music inside the editor.

โœ… Where ChatCut wins
Agent-driven cutting end to end, editable code motion graphics, in-editor generation, and a handoff that opens as a CapCut draft (desktop app only).
โš ๏ธ Where Descript still wins
Sending your edit to other software, since Descript exports many timeline formats and ChatCut exports one. Its word-level voice replacement and studio audio and eye-contact fixes are also more mature, as is team review.
โš ๏ธ The free tiers work in opposite ways, and this is the thing to check before you switch. ChatCut's free exports have no watermark, but the free export allowance is a lifetime total, not a monthly one, so it runs out and never refills. Descript's free tier renews monthly but adds a watermark. If you post several videos a week, budget for a paid plan on either.
read before uploading client work

Rights and your footage ๐Ÿ”’

โœ… Your exports are yours Their terms explicitly don't restrict you creating, exporting and commercially publishing your own output. Sponsored and client work is fine.
โœ… No training on your media Their terms say plainly: "We do not use user media to train AI models." That's a real differentiator and worth crediting.
โš ๏ธ Generated assets are a grey area There's no clause covering who owns AI-generated music or B-roll, and their terms ask you to indemnify them for third-party claims. So if generated material infringes, that lands on you.
โš ๏ธ NDA'd footage needs care Your media isn't trained on, but it still passes through third-party model vendors, and there's no published subprocessor list. Get written client permission, or use their desktop app where footage stays local.
Follow @cindiezhu for more AI you can actually use ๐ŸŒฑ