Cindy Zhu.
← all free guides
Marketing & ads Claude

Bulk ad voiceovers with Claude and Fish Audio

Hey, it's Cindy ๐ŸŒฑ This is the exact setup from the reel, upgraded into the full pipeline: Claude writes ten genuinely different ad scripts, Fish Audio voices them, and then you give each voiceover a video to live in. Test a bunch, run the winner. Studio quality, no studio.

Most ad voiceovers either cost a fortune at a studio or sound like a robot reading a receipt. This skips both. And because a voiceover without a video is only half an ad, this guide now covers the whole thing: script, voice, visuals, and assembly.

the stack

The full stack at a glance ๐Ÿง 

  1. โœ๏ธ Claude writes and runs it. Ten ad scripts with ten different hooks, then it calls the Fish Audio API and generates the voiceovers itself, no extra apps.
  2. ๐ŸŽ™๏ธ Fish Audio is the voice. It performs the emotion you tag inline, so the ads actually sound human.
  3. ๐ŸŽฅ The video layer comes last (or first). Talking-face or b-roll visuals from an AI video tool, and it works in both directions, voice-first or video-first. Step 4 covers both.

ten real hooks

Step 1, the ad-script prompt ๐Ÿ“

This is where most bulk-generated ads die: you ask for ten scripts and get the same ad wearing ten hats. This prompt forces ten genuinely different hooks and a proper short-form structure (hook in the first 3 seconds, one idea per ad, CTA word for word). Paste it into Claude with your details dropped in:

๐Ÿ“ Ten scripts, ten hook types
You are a direct-response copywriter who writes for the ear, not the page. Here is my product and offer: [paste yours here]. My audience: [who it's for]. The one action I want: [e.g. click the link / DM us / buy today].

Write me 10 ad voiceover scripts, each 15 to 30 seconds read aloud (roughly 40 to 75 words). Rules:

1. Structure every script as: hook in the first 3 seconds, the problem in one line, my product as the fix with ONE concrete benefit, one line of proof or believability, then the call to action word for word.
2. Every script opens with a DIFFERENT hook type. Use each of these exactly once: direct callout (name the exact person or pain, like "if you run ads and your costs just doubled..."), bold claim, question, myth-bust, curiosity gap, mini story, stat shock, mistake warning ("stop doing X"), before-and-after, social proof.
3. Make the benefit land like Hormozi's value equation: name the dream outcome, make it believable, and shrink the time and effort ("in minutes, not weeks").
4. Write for the ear: short words, contractions, no jargon, nothing you wouldn't actually say out loud.
5. Mark emotion tags inline exactly where the read should change, like [excited] on the hook, [serious] on the problem, [whisper] on the close. Vary the emotional arc across the ten, they should not all peak in the same place.
6. Label each script with its hook type. If two scripts could swap hooks without anyone noticing, rewrite one.

Output: numbered 1 to 10, spoken words only, no camera directions.

Why the hook rule matters: on short-form, the first three seconds decide everything, and ad fatigue is really hook fatigue. Ten scripts with ten hook types means you're testing ten actual hypotheses, not one ad ten times.


one connection

Step 2, connect Fish Audio to Claude ๐Ÿ”Œ

Claude can call Fish Audio directly through its API, so it generates the audio without you touching another tool.

  1. ๐Ÿ”‘ Make a free account at fish.audio and grab your API key. Put it in a .env file, never paste it into the chat.
  2. ๐Ÿ”Œ Then ask Claude to connect and generate:
๐Ÿ”Œ Generates all ten voiceovers
Connect to the Fish Audio API using the key in my .env file, and never print the key. For each of the 10 ad scripts above, generate a voiceover with the free S2.1 Pro model (model s2.1-pro-free), keep the emotion tags inline so the reads are not flat, and save them as numbered audio files I can listen through.
Heads up: the Fish Audio S2.1 API is free for developers until July 31, so you can test all of this for nothing right now.

the emotion layer

Step 3, make it sound human ๐ŸŽญ

The difference between obviously-AI and wait-that's-a-real-ad is the emotion tags. You type the feeling straight into the script:

๐Ÿ”ฅ [excited] on the hook to grab attention
๐Ÿคซ [whisper] on the close to pull people in

Fish Audio performs the tag instead of reading flat, so every one of your ten reads has real energy.


the visual layer

Step 4, give the voiceover a video to live in ๐ŸŽฅ

A voiceover is half an ad. Here are both directions, pick the one that matches where you're starting from.

๐Ÿ…ฐ Voice-first (you have your ten MP3s):

  • ๐Ÿ—ฃ Want a talking face? Use Open Generative AI, a free open-source studio: open its Lip Sync Studio, upload a portrait (a real photo or an AI-generated presenter) plus your Fish Audio MP3, and it outputs a talking video synced to your read. The app is free; generations run on your own prepaid key at cents per clip.
  • ๐ŸŽฌ Want product shots or cinematic b-roll instead? Generate the visuals in Higgsfield: its Marketing Studio builds product-ad visuals from your product image, and Cinema Studio gives you cinematic shots with real camera controls. Generate one clip per script beat, then cut them under the voiceover in step 5.

๐Ÿ…ฑ Video-first (you already have a visual that works):

๐Ÿ…ฑ Times the script to your cuts
Here are the beats of my video with timestamps: [e.g. 0-3s product close-up, 3-8s hands using it, 8-14s result shot, 14-18s logo]. Write an ad voiceover script that hits each beat at the right moment, using the same hook and emotion tag rules as before. Keep it under [X] seconds total, then generate it in Fish Audio.

Claude times the script to your footage, Fish Audio reads it, and the voiceover lands on your cuts instead of fighting them.

Which order should you use? Voice-first when the message carries the ad (offers, testimonials, explainers). Video-first when the visual is the hook (a product demo, a transformation, anything that stops the scroll on its own).


put it together

Step 5, assemble and ship ๐ŸŽฌ

  1. ๐Ÿ“ฅ Drop the video and the voiceover into CapCut (or your editor), one script per cut.
  2. ๐Ÿ’ฌ Auto-captions on, always. Most feeds play muted until you earn the sound-on.
  3. ๐Ÿ”Š One sound effect per visual reveal, nothing extra.
  4. ๐Ÿ“ฑ Export 9:16 for Reels and TikTok, then test three variants at a time, changing one thing each (usually the hook).

go further

3 bonus prompts to run next ๐ŸŽ

๐ŸŒ Localize a winner.

๐ŸŒ Make a winner sound native
Rewrite my top 3 ad reads for [country or market], adjusting the slang, references, and tone so they sound native, then regenerate them in Fish Audio.

๐ŸŽฏ Scroll-stopping hooks.

๐ŸŽฏ Fifteen more openers
Give me 15 different opening lines for this ad, each under 4 seconds, designed to stop the scroll. Mark the emotion tag for each.

๐Ÿ“Š Pick what to test first.

๐Ÿ“Š Rank before you spend
Based on my product and audience, rank these 10 ad angles by which is most likely to convert, and tell me which 3 to test first and why.

use it well

How to get the most out of it ๐ŸŽ“

๐ŸŒฑ New to this generate just 3 reads first, not 10, and listen back so you can hear the difference the emotion tags make.
๐Ÿ” Already running ads feed Claude your current best ad and ask for 10 variations on that winner, then test them head to head.

what to watch

The honest part ๐Ÿซถ

AI voiceovers are genuinely good now, but they are not magic. The script still has to be a good ad, which is exactly why Claude writes ten different hooks instead of one, so you can test and let the winner emerge. On the video side, budget honestly: Higgsfield is a paid tool, and Open Generative AI charges cents per generation through your own key, so test cheap (one image-based clip) before rendering the expensive version. Always watch the full ad back with sound on and off before you publish, your own eye and ear are the final check.

The links ๐Ÿ”—

๐ŸŽ™ Fish Audio: fish.audio
๐ŸŽฌ Higgsfield (Marketing Studio + Cinema Studio for the visual layer)
๐Ÿ—ฃ Open Generative AI: hosted version ยท official repo (Lip Sync Studio for talking-face ads)
โœ‚๏ธ CapCut: capcut.com (assembly + captions)

Follow @cindiezhu for more AI tips every single day ๐ŸŒฑ