Skip to main content
📖 The AI Tool Bible

Best AI tools for speech to text

24 tools in the Audio category, filtered to speech to text.

All Audio
AudioCraft preview image
AudioCraft logo

AudioCraft

Audio · MusicGen, AudioGen, EnCodec
8.2

Meta's open-source research toolkit for generating music and sound effects from text via a single autoregressive language model.

Free· Free and open source; self-hostedtext-to-musicsound-effects
Respeecher preview image
Respeecher logo

Respeecher

Audio · Proprietary Respeecher voice models
8.1

Studio-grade AI voice cloning and TTS used by Hollywood productions for speech-to-speech and dubbing work.

Freemium· Free trial; TTS API $2/hour pay-as-you-go; custom enterprise pricing for voice cloningvoice-cloningtext-to-speech
Hume AI preview image
Hume AI logo

Hume AI

Audio · Octave, EVI, TADA
8.0

Emotionally intelligent voice AI with expressive TTS, speech-to-speech, and human-feedback evaluation APIs.

Freemium· Free: $0 · Starter: $3 · Creator: $7 · Pro: $70 · Scale: $200expressive-ttsvoice-cloning
Sesame preview image
Sesame logo

Sesame

Audio · Sesame CSM (1B / 3B / 8B)
8.0

Conversational voice AI aiming to cross the uncanny valley with context-aware, emotionally aware speech.

Free· Free research preview; consumer product pricing not announcedconversational-voicetext-to-speech
WellSaid preview image
WellSaid logo

WellSaid

Audio · Proprietary WellSaid TTS
8.0

Enterprise-grade AI text-to-speech built on licensed voice actor recordings.

Freemium· Free trial; paid plans for teams and enterprise (contact sales for API)text-to-speeche-learning narration
Wispr Flow preview image
Wispr Flow logo

Wispr Flow

Audio
8.0

System-wide voice-to-text dictation that auto-edits filler words and learns your jargon.

Freemium· Pro Monthly Plan: $15.00 · Pro Annual Plan: null · Monthly Student Plan: $7.49 · Pro Weekly Plan: $4.49 · Annual Student Plan: nulldictationvoice-to-text
Deepgram preview image
Deepgram logo

Deepgram

Audio · Nova, Flux, Speak (proprietary)
7.9

Production-grade speech-to-text, text-to-speech, and voice-agent APIs for real-time and batch audio.

Freemium· Free credits on signup; usage-based pricing; enterprise contracts availablespeech-to-texttext-to-speech
Voicebox preview image
Voicebox logo

Voicebox

Audio · Multi-model (Chatterbox, Qwen TTS, Whisper, etc.)
7.4

Open-source desktop voice studio for local cloning, dictation, and giving MCP agents a voice.

Free· Free and open source; optional $VOICEBOX token donationsvoice-cloningtext-to-speech
Azure AI Speech (Neural TTS) preview image
Azure AI Speech (Neural TTS) logo

Azure AI Speech (Neural TTS)

Audio · Azure Neural TTS (plus HD and Azure OpenAI voices)
7.3

Microsoft's enterprise-grade neural text-to-speech with 100+ languages, custom brand voices, and SSML control.

Freemium· Free (F0): Free · Pay as You Go: Voice Live Prices: $- · Commitment Tiers – Standard: $- for 2,000 hourstext-to-speechvoice-cloning
Dia preview image
Dia logo

Dia

Audio · Dia-1.6B
7.3

Open-weights 1.6B text-to-dialogue model that generates ultra-realistic multi-speaker conversations in one pass.

Free· Free, open weights (Apache 2.0); hosted larger version waitlisteddialogue-generationvoice-cloning
LOVO AI preview image
LOVO AI logo

LOVO AI

Audio · Proprietary (LOVO Pro V2 voices)
7.2

Text-to-speech and voice cloning platform with 500+ voices, an integrated video editor, and a developer API.

Freemium· 14-day free Pro trial, no credit card; paid subscription tierstext-to-speechvoice-cloning
Veritone Voice preview image
Veritone Voice logo

Veritone Voice

Audio · Proprietary (Veritone aiWARE)
7.2

Enterprise-grade voice cloning and synthesis platform built for broadcasters, studios, and large media operations.

Enterprise· Contact sales / demo onlyvoice-cloningtext-to-speech
AI Song Maker preview image
AI Song Maker logo

AI Song Maker

Audio · Multi-model (ACE-Step, MusicGen, DiffRhythm, Riffusion)
7.1

Browser-based song generator that wraps multiple open music models behind a single freemium UI.

Freemium· Free: $0 · Basic: $14.99 · Standard: $29.99 · Pro: $59.99text-to-songlyrics-generation
iSpeech preview image
iSpeech logo

iSpeech

Audio
7.0

Veteran cloud TTS and speech recognition API with broad SDK and language coverage.

Freemium· Free mobile SDK for non-revenue apps; ~$0.0001-$0.05 per word/transactiontext-to-speechspeech-recognition
Loudly preview image
Loudly logo

Loudly

Audio · Proprietary Loudly AI
7.0

AI music generator with royalty-free output, stem splitting, and distribution to Spotify and friends.

Freemium· Free tier; paid plans on /music/pricingtext-to-musicroyalty-free background music
MockingBird preview image
MockingBird logo

MockingBird

Audio · GE2E + Tacotron + HiFi-GAN/WaveRNN/Fre-GAN
7.0

Open-source Mandarin-first voice cloning that mimics a speaker from a 5-second sample.

Free· Free, open source (MIT)voice-cloningtext-to-speech
Remusic preview image
Remusic logo

Remusic

Audio · Remusic V4 Pro (proprietary)
7.0

All-in-one AI music studio that bundles song generation, voice cloning, stem splitting, and karaoke tools.

Freemium· Pro: $20.9/mo · Basic: $7.9/mo · Starter: $4.1/motext-to-musicvoice-cloning
ZenMic preview image
ZenMic logo

ZenMic

Audio
7.0

Text-to-podcast generator with multi-speaker AI voices and RSS publishing.

Freemium· Monthly: $19 · Yearly: ≈ $8.25/mo · Early Adopter Tier: ?text-to-podcastcontent-repurposing
Murf AI preview image
Murf AI logo

Murf AI

Audio · Murf Gen2 / Murf Falcon
6.9

Studio-grade text-to-speech and real-time voice agents with 200+ voices across 35+ languages.

Freemium· Free: $0 / month · Creator: $19 / month · Business: $66 / month · Enterprise: Customtext-to-speechvoice-cloning
WhisperAPI preview image
WhisperAPI logo

WhisperAPI

Audio · OpenAI Whisper
6.9

Hosted OpenAI Whisper transcription with a pay-as-you-go API and drop-in web dashboard.

Paid· 20 API Credits: $5 · 100 API Credits: $20 · 200 API Credits: $30 · Custom: $100.00audio-transcriptionvideo-subtitles
Audify AI preview image
Audify AI logo

Audify AI

Audio · OpenAI TTS (tts-1, tts-1-hd, gpt-4o-mini-tts)
6.8

Pay-as-you-go web wrapper around OpenAI's text-to-speech voices.

Freemium· BYO OpenAI key (free); or top up from $2 pay-per-usetext-to-speechvoiceover
Fish Audio preview image
Fish Audio logo

Fish Audio

Audio · Fish Audio S2.1 Pro (in-house); S1 and S2 checkpoints open-sourced

Expressive, emotion-controllable text-to-speech and voice cloning with an open-model heritage

Freemium· Basic: $10 · Pro: $20 · Enterprise: Contact salesYouTube video voiceoverAudiobook narration
Horch preview image
Horch logo

Horch

Audio · Whisper (local) for transcription; optional Ollama / MLX local LLMs for summarization

Privacy-first, on-device meeting assistant for macOS

Paid· One-time purchase: €49 onceConfidential client meeting transcriptionAutomatic action-item extraction
Kyutai Moshi preview image
Kyutai Moshi logo

Kyutai Moshi

Audio · Moshi (7B-class speech-text foundation model) + Mimi neural audio codec, in-house by Kyutai

Open-source, full-duplex speech-to-speech foundation model with sub-200ms latency

Free· Free and open source. Models under CC-BY 4.0, code under MIT (Python) / Apache 2.0 (Rust). Self-hosted only — you pay your own compute (24GB+ GPU for PyTorch, or Apple Silicon via MLX).Real-time voice assistant prototypesResearch on full-duplex spoken dialogue