Audio
Voice cloning, music generation, speech-to-text.
60 tools
AI audio has split cleanly into three lanes: speech synthesis (TTS + voice cloning), music generation, and speech-to-text — each with a clear leader.
Covers voice cloning and TTS (ElevenLabs, Resemble.ai, Murf), AI music generation (Suno, Udio), and speech-to-text (AssemblyAI, Whisper).
Pick ElevenLabs for voice quality. Pick Suno or Udio for AI music. Pick AssemblyAI when you need diarisation and timestamps; pick Whisper when you can self-host and want zero cost.

Otter.ai
AI meeting notetaker that transcribes calls, summarizes them, and pulls out action items in real time.

Soundful
Template-driven AI music generator that spits out royalty-free, commercially licensable tracks in seconds.

Scribbl
Bot-free AI meeting recorder, transcriber, and summarizer for Google Meet.

Bland AI
Enterprise voice AI for automated phone calls at scale

Fish Audio
Expressive, emotion-controllable text-to-speech and voice cloning with an open-model heritage

Horch
Privacy-first, on-device meeting assistant for macOS

Kyutai Moshi
Open-source, full-duplex speech-to-speech foundation model with sub-200ms latency

Retell AI
Build, test and deploy production-grade AI voice agents for inbound and outbound calls.

so-vits-svc
SoftVC VITS Singing Voice Conversion — open-source pipeline for training and running singing-voice models.

Threadfork
Private, local-first AI meeting notetaker for macOS

Vapi
Developer platform for building, deploying, and scaling production voice AI agents

VoxAI
Local macOS conversation recorder with live AI copilot and speaker-labelled transcripts