AI Tools

Audio & Voice Tools

12 tools

AIVA

#6
Audio & Voice Tools

AIVA is "the AI composer that hands back an editable score, not just a rendered track." No lyrics, no vocals: you choose from 250+ style presets (Modern Cinematic, Symphonic, Chinese, Jazz, Ambient, Electronic, Tango…) or upload a reference MIDI/audio to train a custom style model, set key/tempo/length/emotion, and it composes a full multi-section arrangement. The built-in editor shows piano-roll score — move a note, change velocity, re-voice the brass, swap instruments — then export MIDI to finish in Logic/Cubase/Pro Tools, or WAV+STEMS on Pro for direct-to-media. 2017 it became the first AI registered as a composer by SACEM, which is why its copyright ladder (Free→Standard→Pro) reads like a music-publishing contract, not a SaaS EULA. Pitch: Suno writes the sung single; AIVA drafts the film/game score you own and can re-orchestrate by hand.

Details
Freemium

Auphonic

#10
Audio & Voice Tools

Auphonic (Auphonic GmbH, Austria 2012/2013) Status: Active (founded 2012/2013 by Hans-Jürgen Schüler et al., Graz; 2 hrs free/mo permanent; 2026-01 RMS-based loudness for Audible/ACX, 2025-12 Denoising Editor; auphonic.com + API + CLI live 2026-09). Category ai-productivity-data; pipeline_stage=none. Not a podcast editor (→ Descript), not a live call denoiser (→ Krisp)​ — it is the broadcast-standard…

Details
Freemium

Cleanvoice

#22
Audio & Voice Tools

Cleanvoice is a freemium AI audio/video cleanup web app that automatically removes filler words, mouth sounds, dead air, and background noise from podcasts, then exports transcripts, summaries, and DAW-compatible timeline markers. Priced by processed audio hour (30 min free trial, ~$1.10/hr on entry subscription); it is a post-production cleanup pass, not a text editor, DAW, or echo remover.

Details
Freemium
Descript

Descript

#41
Audio & Voice Tools

Descript (San Francisco, 2017; Andrew Mason ex-Google; Series C ~$100M, backers a16z/GV) = transcript-first AI audio/video editor​ — you edit by deleting/retyping words in a transcript and the waveform + video cut follow, on top of which sits Underlord (agentic AI co-editor), Studio Sound (one-pass denoise), Overdub (voice clone fix), Eye Contact (gaze correction), Regenerate (retype mis-spoken line → mouth + voice resynced), 30+ lang translate/dub, AI avatar + Veo3.1/Nano Banana video gen. Not Premiere/Final Cut (timeline NLE), not Riverside (remote capture-first), not Opus Clip/Captions (clip-only), not Otter (notes-only). Your 42nd CPT card → ai-productivity-data; pipeline_stage none (spans transcribe → text-edit → AI-clean → clip → dub → export). Table-first, Claude Code skeleton. 2026 trap: AI credits not hours are the real meter​ — Studio Sound ~10 credits/min, Hobbyist 400/mo dies on a 40-min ep, and formerly-"unlimited" features (Overdub regen, eye contact) are now credit-metered.

Details
ElevenLabs

ElevenLabs

#46
Audio & Voice Tools

Short Description: AI audio platform spanning text-to-speech (Eleven v3), instant/professional voice cloning, multilingual dubbing (Dubbing v2), Scribe v2 STT, Eleven Music v2, sound effects, and ElevenAgents conversational voice agents — credit-metered across all products, 70+ languages, 5,000+ voices.

Details

Krisp

#81
Audio & Voice Tools

Krisp = an on-device real-time voice-clarity layer that sits between your mic/speaker and any call app (Zoom, Meet, Teams, Discord, OBS, DAW), stripping background noise, echo, and other people's voices bidirectionally during live calls, then optionally turning the meeting into bot-free transcript + summary + action items. It is a live capture shield, not a post-production cleanup tool.

Details

Otter.ai

#110
Audio & Voice Tools

Otter.ai = a cloud-based AI meeting notetaker that joins (or watches) Zoom / Google Meet / Teams via OtterPilot, streams live captions, speaker-labels the transcript, then emits summary + action items + searchable archive + CRM sync. It captures and documents speech; it does not generate voice (Murf), cancel noise (Krisp), or remove filler words (Cleanvoice).

Details

Resemble AI

#129
Audio & Voice Tools

Resemble AI = enterprise-first generative voice platform (TTS + speech-to-speech + 10-sec Rapid Clone + Professional Clone) that ships with its own trust stack — PerTh watermarking on every render and Resemble Detect (DETECT-3B / DETECT-World) for audio/image/video deepfake screening — billed per synthesized second on Flex, or custom on Enterprise with on-prem/air-gapped deploy. It is the security-and-compliance choice among voice generators, not a consumer narrator like Murf/ElevenLabs.

Details

Suno

#144
Audio & Voice Tools

Suno = prompt-to-full-song AI music generator (vocals + lyrics + full arrangement) with the v5.5 model, verified vocal cloning (Voices), custom-style fine-tunes (Custom Models), Suno Studio browser DAW, and stem export — credit-based freemium, no public API, downloads capped from 2026-09-03. It is the highest-volume consumer AI music app (~7M tracks/day), not a voice generator or an ASR tool.

Details

Udio

#155 IOS only
Audio & Voice Tools

Udio = prompt-to-song AI music generator (v1.5 / Allegro v1.5) with section-level inpainting, style blending, voice control, and 48 kHz-class stereo output, but since the Oct 2025 UMG + Warner settlements all audio/video/stem downloads are disabled on every tier — output stays in-platform under a licensed "walled garden". Credit-metered freemium, iOS app only, no public API. It is the quality/control-first cousin of Suno with the strictest export lock in the category.

Details

Voicemod

#161
Audio & Voice Tools

Voicemod is a real-time AI voice changer and soundboard that sits between your microphone and any app accepting mic input (Discord, OBS, Zoom, Google Meet, Fortnite, Valorant, VRChat, Twitch). It installs a virtual audio device ("Voicemod Virtual Microphone"), applies speech-to-speech voice transformation with ultra-low local latency (<20 ms claimed), and pipes the modulated signal out as if it were a physical mic. Library spans 200+ first-party voices (Robot, Alien, Demon, anime, celeb-style, emotional) plus 300,000+ community voices and 800,000+ community soundboard clips. VoiceLab (Pro) lets users stack reverb/delay/vocoder/pitch/formant/Robotifier into saved custom voices. AI voices are Fairly Trained-certified, trained on pro voice-actor recordings. On Snapdragon X Elite, AI workload offloads to NPU. Not a TTS engine (no text-in-audio-out), not a DAW plugin, not an audiobook narrator — it modulates a live human voice as you speak.

Details

Whisper

#164 OpenAI
Audio & Voice Tools

Not shut down.​ OpenAI Whisper launched Sep 2022 (MIT-licensed weights on GitHub), and as of 2026-08 the open-source repo is intact, whisper-1 is still callable on the OpenAI API at $0.006/min, and OpenAI has added GPT-4o Transcribe / Mini Transcribe / Diarize as newer hosted siblings on the same /v1/audio/transcriptions path — Whisper itself was not renamed, merged, or retired. Category ai-productivity-data; pipeline_stage none. It is a model + reference runtime, not a consumer app: no dashboard, no login, no subscription from OpenAI. Killer differentiator = only mainstream ASR you can run fully offline under MIT with 99-language coverage; not Deepgram (real-time-first), not AssemblyAI (managed intelligence layer), not Otter/Descript (end-user products), not gpt-4o-transcribe (higher accuracy, no SRT/VTT). 2026 trap: base Whisper hallucinates on silence, no native diarization/real-time, 25MB/25-min API cap, large-v3 needs ~10GB VRAM.

Details

← View all AI tools