Skip to main content
📖 The AI Tool Bible

Audio

Voice cloning, music generation, speech-to-text.

60 tools

Why it matters

AI audio has split cleanly into three lanes: speech synthesis (TTS + voice cloning), music generation, and speech-to-text — each with a clear leader.

What's in here

Covers voice cloning and TTS (ElevenLabs, Resemble.ai, Murf), AI music generation (Suno, Udio), and speech-to-text (AssemblyAI, Whisper).

How to pick

Pick ElevenLabs for voice quality. Pick Suno or Udio for AI music. Pick AssemblyAI when you need diarisation and timestamps; pick Whisper when you can self-host and want zero cost.

Otter.ai preview image
Otter.ai logo

Otter.ai

Audio · Proprietary speech + LLM stack
6.8

AI meeting notetaker that transcribes calls, summarizes them, and pulls out action items in real time.

Freemium· Basic: $20 · Pro: $50 · Enterprise: Contact salesmeeting-transcriptionmeeting-summaries
Soundful preview image
Soundful logo

Soundful

Audio · Proprietary (human-aided AI)
6.8

Template-driven AI music generator that spits out royalty-free, commercially licensable tracks in seconds.

Freemium· Free tier; Plus/Pro/Business monthly per-user; Enterprise on requestbackground-musiccontent-creator-audio
Scribbl preview image
Scribbl logo

Scribbl

Audio
6.7

Bot-free AI meeting recorder, transcriber, and summarizer for Google Meet.

Freemium· Free: 15 meetings/month; paid team plans for shared libraries and CRM integrationsmeeting-transcriptionmeeting-summaries
Bland AI preview image
Bland AI logo

Bland AI

Audio · Proprietary in-house voice models

Enterprise voice AI for automated phone calls at scale

Enterprise· Start: $0 · Build: $299 · Scale: $499 · Enterprise: CustomOutbound appointment remindersInsurance claims intake calls
Fish Audio preview image
Fish Audio logo

Fish Audio

Audio · Fish Audio S2.1 Pro (in-house); S1 and S2 checkpoints open-sourced

Expressive, emotion-controllable text-to-speech and voice cloning with an open-model heritage

Freemium· Basic: $10 · Pro: $20 · Enterprise: Contact salesYouTube video voiceoverAudiobook narration
Horch preview image
Horch logo

Horch

Audio · Whisper (local) for transcription; optional Ollama / MLX local LLMs for summarization

Privacy-first, on-device meeting assistant for macOS

Paid· One-time purchase: €49 onceConfidential client meeting transcriptionAutomatic action-item extraction
Kyutai Moshi preview image
Kyutai Moshi logo

Kyutai Moshi

Audio · Moshi (7B-class speech-text foundation model) + Mimi neural audio codec, in-house by Kyutai

Open-source, full-duplex speech-to-speech foundation model with sub-200ms latency

Free· Free and open source. Models under CC-BY 4.0, code under MIT (Python) / Apache 2.0 (Rust). Self-hosted only — you pay your own compute (24GB+ GPU for PyTorch, or Apple Silicon via MLX).Real-time voice assistant prototypesResearch on full-duplex spoken dialogue
Retell AI preview image
Retell AI logo

Retell AI

Audio · LLM-agnostic (OpenAI, Anthropic, Google and others selectable per agent); proprietary turn-taking model and voice pipeline in-house

Build, test and deploy production-grade AI voice agents for inbound and outbound calls.

Freemium· Pay-as-you-go: $0.07-$0.31/minute · Enterprise: Custom PricingAI inbound call answeringOutbound lead qualification
so-vits-svc preview image
so-vits-svc logo

so-vits-svc

Audio · SoftVC content encoder + VITS backbone + NSF-HiFiGAN vocoder; optional ContentVec, HuBERT-Soft, Whisper-PPG, WavLM encoders and shallow-diffusion module.

SoftVC VITS Singing Voice Conversion — open-source pipeline for training and running singing-voice models.

Free· Free / open-source (AGPL-3.0). You provide your own compute (typically a CUDA-capable GPU) and training datasets.Singing voice conversion (AI covers)VTuber and virtual-character singing voices
Threadfork preview image
Threadfork logo

Threadfork

Audio · On-device local LLMs running on Apple Silicon (specific model family undisclosed)

Private, local-first AI meeting notetaker for macOS

Paid· Pro: $39Confidential client meeting transcriptionSales call notes and follow-ups
Vapi preview image
Vapi logo

Vapi

Audio · Model-agnostic: OpenAI GPT-4o, Anthropic Claude, Google Gemini, Groq, DeepSeek, Llama; STT/TTS via Deepgram, ElevenLabs, PlayHT, Cartesia, Azure

Developer platform for building, deploying, and scaling production voice AI agents

Freemium· Build: Usage based · Scale: Contact UsAI phone receptionistOutbound lead qualification calls
VoxAI preview image
VoxAI logo

VoxAI

Audio · Bring-your-own via MCP (Claude, GPT, or other MCP-compatible clients)

Local macOS conversation recorder with live AI copilot and speaker-labelled transcripts

Paid· FREE TRIAL: US$0 · EARLY BIRD · REGULAR: US$19.90Real-time meeting transcription with speaker labelsLive AI copilot during sales calls and negotiations