Cindy Zhu.
← all free guides
Content creation Multi-tool

Edit a whole video with AI

Hey, it's Cindy ๐ŸŒฑ You commented OMNI, so here is the whole thing: the prompt, the click path, and the settings that actually matter. The reel you watched was one take on a tripod. Every camera move, motion graphic and text effect in it was added afterwards by Arcads running Google's Omni 1.1. No timeline, no CapCut. ๐Ÿ‘ฉโ€๐Ÿ’ป
The mindset shift, and it is the whole thing. You are not editing any more, you are directing. In a timeline you assemble what you filmed. Here you describe the shot you wish you had filmed, and it makes it from the footage you have. So the skill that matters stops being "can I use the software" and becomes "can I describe a shot properly". That is why the prompt below is written the way it is.

the setup

The click path ๐Ÿ› ๏ธ

  1. Shoot one clean take. Tripod, locked off, don't move. Resist the urge to add movement in camera, because a static take gives the model the most room to invent a move. Good light and a clean background do more for the result than anything you type.
  2. New Project, then Video. In Arcads, start a new project and choose Video.
  3. Pick the Omni 1.1 Flash model. This is the one that does the camera work and the on-screen elements. If you pick a different model the prompt below will not behave the same way.
  4. Draft at 360p. Do not generate your first attempt at full resolution. Drafts at 360p cost roughly a third, and you are going to throw most of them away, because the first version of any shot is a guess.
  5. Paste your clip in with the prompt, and generate three versions. Not one. Camera moves vary between runs and you want to pick, not accept.
  6. Upscale only the keeper to 4K. This is where the money goes, so it goes on exactly one file.
๐ŸŽฅ the prompt
Take this locked-off tripod clip and edit it as a finished short-form video, keeping my performance and audio exactly as they are.

Camera: build movement that was never filmed. Start on a slow push in, then pull the camera back to reveal the room, then orbit around me and settle slightly above eye line. Keep every move slow and motivated, the kind a real operator would make, never fast or swooping. My face stays sharp and centred through every move.

On-screen: add motion graphics and text that support what I am actually saying, appearing on the words they relate to and clearing before the next point. Keep the type clean and readable on a phone, and keep it out of the lower third where captions sit.

Look: keep the grade natural and consistent with the original footage. Do not smooth my skin, do not change my face, do not restyle the room.

Cuts: hold each angle long enough to read, roughly two to four seconds, and change angle on a natural pause in my speech rather than mid-sentence.
๐ŸŒฑ Change one thing at a time. When a result is nearly right, resist rewriting the whole prompt. Change only the sentence covering the thing you disliked and regenerate. Rewriting everything gives you a completely different video and you lose the parts that were already working.

The two features that change how you work ๐Ÿง 

Ten seconds of memory, not one
When a clip runs short you tell it to keep going, and because it now holds around ten seconds of context instead of one, the continuation still matches. A shot can run out to about forty seconds and hold together. Practically: you no longer need to reshoot because a take was two seconds shy.
Draft cheap, upscale once
360p drafts cost roughly a third. The workflow that follows from that is generate a lot, judge, then spend. Most people do the opposite, generate one expensive version and then talk themselves into liking it.

Five prompt lines worth stealing ๐Ÿ“‹

Add any of these to the prompt above. Each one fixes a specific failure I hit.

๐ŸŽฌ fixes
the move is too fast and swoopy
Slow every camera move down by half. The camera should drift rather than travel. Think a tripod head being panned gently by hand, not a drone.

it keeps changing my face
Treat my face as locked reference. Do not smooth, slim, relight or alter my features in any frame. If a camera move would require inventing detail on my face, choose a smaller move instead.

the text covers my captions
Keep all on-screen type in the upper two thirds of the frame. The lower third is reserved and must stay clear.

it cuts mid-sentence
Change angle only on a natural pause or breath in my speech. Never cut across a word. If there is no pause, hold the angle longer.

the graphics feel generic
The on-screen elements should illustrate the specific thing I am saying at that moment, not float as decoration. If a line names an object, show that object. If a line names a number, put that number on screen.

The honest bit โœ…

The take still decides the ceiling
This adds movement and polish to what you filmed. It does not fix a mumbled delivery, a badly lit room or a weak script. One clean take in good light beats any prompt.
Expect to regenerate
Three drafts to get a keeper is normal. That is exactly why you draft at 360p, and why "generate one and hope" is the expensive way to work.
Check the hands and the edges
Generated motion is strongest on faces and weakest at frame edges and on hands. Watch a draft full screen once before you spend the upscale.
It is credits, not free
Every generation costs, drafts included. The 360p-first workflow is what keeps a session affordable rather than a surprise.

The links ๐Ÿ”—

๐ŸŽฌ Arcads: arcads.ai

Follow @cindiezhu for more AI you can actually use ๐ŸŒฑ