Skip to main content
📖 The AI Tool Bible

AI tools tagged Supports Audio

29 tools matching this tag.model

All tags →
EL

ElevenLabs

Featured
Audio · ElevenLabs Multilingual v2
9.4

The gold standard for AI voice cloning and TTS.

Freemium· Free: $0 · Starter: $6 · Creator: $11 · Pro: $99 · Scale: $299TTSvoice cloning
G4

GPT-4o

Featured
Writing · GPT-4o
9.4

OpenAI's multimodal flagship behind ChatGPT.

Freemium· Basic: $10 · Pro: $30 · Enterprise: Contact salesgeneral writingsummarization
SU

Suno

Featured
Audio · Suno v4
9.2

Text-to-song AI — full vocal tracks from a prompt.

Freemium· Free Plan: $0 · Pro Plan: $8 · Premier Plan: $24songwritingdemos
UD

Udio

Audio · Udio (proprietary)
8.8

Suno's main rival for AI-generated full songs.

Freemium· Free; Standard $10/mo; Pro $30/mofull songsmusic demos
WH

Whisper

Audio · Whisper large-v3
8.6

OpenAI's open-source speech-to-text — the de-facto baseline.

Free· Free open weights; $0.006/min via OpenAI APItranscriptionself-hosted
MU

Mubert

Audio · Proprietary sample-based generative engine
8.3

AI music generator that spits out royalty-free background tracks for video, podcast, and app use.

Freemium· Ambassador: Free · Creator: $14 · Pro Most popular: $39 · Business: $199background-musicroyalty-free-soundtracks
SO

Soundraw

Audio · Proprietary in-house model
8.2

AI music generator that spits out royalty-free, customizable tracks by genre and mood.

Freemium· Creator: €5.83/mo · Artist Starter: €10.75/mo · Artist Pro: €12.42/mo · Artist Unlimited: €17.42/mo · Enterprise: Askbackground musicvideo soundtracks
AI

AIVA

Audio · Proprietary (AIVA)
8.1

AI music composition tool that generates royalty-friendly tracks in 250+ styles with editable MIDI output.

Freemium· Free; Standard ~€11/mo, Pro ~€33/mo (billed yearly)music-generationsoundtrack-composition
LO

LocalAI

Writing · Multi-model (llama.cpp, diffusers, whisper, etc.)
8.1

Self-hosted OpenAI-compatible API for running LLMs, image, and audio models on your own hardware.

Free· Free and open source (MIT)local-llm-inferenceopenai-api-replacement
HA

Hume AI

Audio · Octave, EVI, TADA
8.0

Emotionally intelligent voice AI with expressive TTS, speech-to-speech, and human-feedback evaluation APIs.

Freemium· Free: $0 · Starter: $3 · Creator: $7 · Pro: $70 · Scale: $200expressive-ttsvoice-cloning
MI

MiniMax

Agents · MiniMax M3, Hailuo 2.3, Speech 2.8, Music 2.6
8.0

Chinese frontier-model lab shipping multimodal foundation models with a 1M-context coding/agent stack.

Freemium· Free tier; Token plan from ~$20/mo (~12.5B tokens); enterprise pricing on requestcoding-agentlong-context
BO

Boomy

Audio · Proprietary (undisclosed)
7.9

Generative AI music maker that lets anyone produce a song in under a minute and push it to Spotify.

Freemium· Free: $0.00 · Recommended: $14.99 · Pro: $39.99ai-music-generationsong-creation
DI

Dia

Audio · Dia-1.6B
7.3

Open-weights 1.6B text-to-dialogue model that generates ultra-realistic multi-speaker conversations in one pass.

Free· Free, open weights (Apache 2.0); hosted larger version waitlisteddialogue-generationvoice-cloning
AS

AI Song Maker

Audio · Multi-model (ACE-Step, MusicGen, DiffRhythm, Riffusion)
7.1

Browser-based song generator that wraps multiple open music models behind a single freemium UI.

Freemium· Free: $0 · Basic: $14.99 · Standard: $29.99 · Pro: $59.99text-to-songlyrics-generation
SA

Stable Audio

Audio · Stable Audio 3.0 (Large/Medium/Small/Small SFX)
7.1

Stability AI's generative audio model family for music and sound effects, with open weights for the smaller variants.

Freemium· Free web app tier; API metered; enterprise licensing for Large modelmusic-generationsound-effects
BA

Beatoven.ai

Audio · Maestro (proprietary)
7.0

Text-to-music generator that spits out royalty-free background tracks with a clean licensing story.

Freemium· Basic: $10 · Pro: $20 · Enterprise: Contact salesbackground-musicsound-effects
IS

iSpeech

Audio
7.0

Veteran cloud TTS and speech recognition API with broad SDK and language coverage.

Freemium· Free mobile SDK for non-revenue apps; ~$0.0001-$0.05 per word/transactiontext-to-speechspeech-recognition
ZE

ZenMic

Audio
7.0

Text-to-podcast generator with multi-speaker AI voices and RSS publishing.

Freemium· Monthly: $19 · Yearly: ≈ $8.25/mo · Early Adopter Tier: ?text-to-podcastcontent-repurposing
ME

Melies

Video · Multi-model (Flux, Kling, Runway, Luma, Hailuo, ElevenLabs, GPT, Claude)
6.9

All-in-one AI filmmaking studio that stitches image, video, voice, and music models into a single production pipeline.

Freemium· Starter: $9/month · Pro: $49/month · Max: $99/monthai-filmmakingtext-to-video
MA

Murf AI

Audio · Murf Gen2 / Murf Falcon
6.9

Studio-grade text-to-speech and real-time voice agents with 200+ voices across 35+ languages.

Freemium· Free: $0 / month · Creator: $19 / month · Business: $66 / month · Enterprise: Customtext-to-speechvoice-cloning
SB

Sera by Twig

Agents
6.9

AI front-desk agent that answers calls, chats, and emails 24/7 and books the appointment.

Freemium· Free: $0 · Starter: $99 · Growth: $499 · Scale: $1,499ai-receptionistappointment-booking
WH

WhisperAPI

Audio · OpenAI Whisper
6.9

Hosted OpenAI Whisper transcription with a pay-as-you-go API and drop-in web dashboard.

Paid· 20 API Credits: $5 · 100 API Credits: $20 · 200 API Credits: $30 · Custom: $100.00audio-transcriptionvideo-subtitles
AI

AInterview

Audio
6.8

AI host that interviews you and turns the conversation into a finished podcast.

Freemium· Free 10 min/mo; Premium $19/mo (2 hrs); pay-as-you-go $0.13-$0.20/minai-podcast-hostsolo-podcasting
AA

Audify AI

Audio · OpenAI TTS (tts-1, tts-1-hd, gpt-4o-mini-tts)
6.8

Pay-as-you-go web wrapper around OpenAI's text-to-speech voices.

Freemium· BYO OpenAI key (free); or top up from $2 pay-per-usetext-to-speechvoiceover
EM

Ecrett Music

Audio
6.8

AI background-music generator that spits out royalty-free instrumental tracks by scene, mood, and genre.

Freemium· Free preview tier; Individual $4.99/mo annual ($7.99 monthly); Business $14.99/mo annual ($24.99 monthly)background-musicyoutube-soundtracks
S2

Seedance 2.0

Video · Seedance 2.0
6.8

ByteDance's multimodal video model with joint audio-video generation and director-level camera control.

Paid· Not disclosed on the page; API metered via ByteDance's platformtext-to-videoimage-to-video
SH

ShortVideoGen

Video · OpenAI Sora 2
6.8

Text-to-video generator wrapping OpenAI's Sora 2 with audio output and per-video credit pricing.

Freemium· 5 free credits; Basic $8.99/mo (50 videos), Standard $14.99/mo (100), Pro $34.99/mo (500)text-to-videoshort-form video
KM

Kyutai Moshi

Audio · Moshi (7B-class speech-text foundation model) + Mimi neural audio codec, in-house by Kyutai

Open-source, full-duplex speech-to-speech foundation model with sub-200ms latency

Free· Free and open source. Models under CC-BY 4.0, code under MIT (Python) / Apache 2.0 (Rust). Self-hosted only — you pay your own compute (24GB+ GPU for PyTorch, or Apple Silicon via MLX).Real-time voice assistant prototypesResearch on full-duplex spoken dialogue
LV

LTX Video

Video · LTX-Video (LTX-2), in-house DiT — 13B full, 13B distilled, 2B distilled, FP8 quantized variants

Open-source DiT video model with synchronized audio, 4K output, and multi-keyframe control

Freemium· Model weights free under OpenRail-M (self-host). Hosted access via LTX Studio (free tier + paid plans), Fal.ai and Replicate pay-per-generation (typically fractions of a cent per second of video). Enterprise licensing available from Lightricks.Text-to-video generationImage-to-video animation