Disclosure: Some links on this page are affiliate links. We may earn a commission if you purchase through them — at no extra cost to you. Read our full affiliate disclosure.

Ai workflow

Clone & Generate Audio

Clone a voice or generate speech, music, and sound — then master and publish.

6 stages 17 tools 6 prompts
Experience
Budget
6 Stages, with a clear route
17 Hand-picked tools
6 Prompt templates to copy
Published audio What you walk away with

Core tools in these workflows

AI Audio/Video Mashup: Suno Udio Runway Descript
Batch Short-video Clipping: Whisper ChatGPT FFmpeg Descript
Article to Explainer Video: ChatGPT ElevenLabs Runway Gamma
Step flow

How the work actually flows

Read the goal → pick a tool → copy the prompt → ship the output

1

Define the Audio

Choose what you need: voice clone, TTS, music, or sound effects, and the target length.

Goal

A brief that names what you actually need.

Prompt template
I need audio for [use case], length [X]. Tell me whether voice cloning, TTS, music, or SFX fits best and what I must prepare first.
Output Brief
Tip Know your usage rights before cloning any voice.
2

Clone the Voice

Upload clean consent-based samples to build a voice model.

Goal

A consented voice model that sounds like the speaker.

Prompt template
Checklist for cloning [voice]: sample length, recording conditions, consent wording, and the checks that tell me the model is ready.
Output Voice model
Tip Use 1–3 min of noise-free speech for best fidelity.
3

Generate Speech / Music

Synthesize the narration or compose the track from a prompt or lyrics.

Goal

Usable narration or music generated from a prompt.

Prompt template
Generate [narration / music] for [project]. Tone: [tone]. Pace: [slow / normal / fast]. Give the prompt plus the pacing or SSML marks to add.
Output Audio (wav / mp3)
Tip Add punctuation and pauses to control pacing.
4

Edit & Clean

Trim, remove breaths, and stitch takes together.

Goal

Clean, evenly leveled audio with no filler.

Prompt template
Clean-up plan for this audio: breaths, mouth clicks, long silences, and the loudness target I should normalize to. List the steps in order.
Output Edited audio
Tip Normalize loudness so episodes sound consistent.
5

Master

Apply EQ, compression, and limiting for a polished, platform-ready sound.

Goal

Platform-ready loudness and tone.

Prompt template
Mastering settings for [podcast / music / ad] on [platform]: LUFS target, EQ moves, compression ratio, and true peak ceiling.
Output Mastered audio
Tip Target -14 LUFS for streaming.
6

Publish

Export and upload to podcast / music / video platforms.

Goal

The audio live, with metadata and raw stems kept.

Prompt template
Write show notes, chapter timestamps, and a title for this episode, then list the export settings for [platform].
Output Published audio
Tip Keep the raw stems for future remixes.

Browse the workflows under Video & Audio Production — these are the source cards feeding this builder.

View all Video & Audio Production workflows →

Do Not Sell or Share My Personal Information Cookie Settings