AI tools tagged Supports Audio
29 tools matching this tag.model
ElevenLabs
FeaturedThe gold standard for AI voice cloning and TTS.
GPT-4o
FeaturedOpenAI's multimodal flagship behind ChatGPT.
Suno
FeaturedText-to-song AI — full vocal tracks from a prompt.
Udio
Suno's main rival for AI-generated full songs.
Whisper
OpenAI's open-source speech-to-text — the de-facto baseline.
Mubert
AI music generator that spits out royalty-free background tracks for video, podcast, and app use.
Soundraw
AI music generator that spits out royalty-free, customizable tracks by genre and mood.
AIVA
AI music composition tool that generates royalty-friendly tracks in 250+ styles with editable MIDI output.
LocalAI
Self-hosted OpenAI-compatible API for running LLMs, image, and audio models on your own hardware.
Hume AI
Emotionally intelligent voice AI with expressive TTS, speech-to-speech, and human-feedback evaluation APIs.
MiniMax
Chinese frontier-model lab shipping multimodal foundation models with a 1M-context coding/agent stack.
Boomy
Generative AI music maker that lets anyone produce a song in under a minute and push it to Spotify.
Dia
Open-weights 1.6B text-to-dialogue model that generates ultra-realistic multi-speaker conversations in one pass.
AI Song Maker
Browser-based song generator that wraps multiple open music models behind a single freemium UI.
Stable Audio
Stability AI's generative audio model family for music and sound effects, with open weights for the smaller variants.
Beatoven.ai
Text-to-music generator that spits out royalty-free background tracks with a clean licensing story.
iSpeech
Veteran cloud TTS and speech recognition API with broad SDK and language coverage.
ZenMic
Text-to-podcast generator with multi-speaker AI voices and RSS publishing.
Melies
All-in-one AI filmmaking studio that stitches image, video, voice, and music models into a single production pipeline.
Murf AI
Studio-grade text-to-speech and real-time voice agents with 200+ voices across 35+ languages.
Sera by Twig
AI front-desk agent that answers calls, chats, and emails 24/7 and books the appointment.
WhisperAPI
Hosted OpenAI Whisper transcription with a pay-as-you-go API and drop-in web dashboard.
AInterview
AI host that interviews you and turns the conversation into a finished podcast.
Audify AI
Pay-as-you-go web wrapper around OpenAI's text-to-speech voices.
Ecrett Music
AI background-music generator that spits out royalty-free instrumental tracks by scene, mood, and genre.
Seedance 2.0
ByteDance's multimodal video model with joint audio-video generation and director-level camera control.
ShortVideoGen
Text-to-video generator wrapping OpenAI's Sora 2 with audio output and per-video credit pricing.
Kyutai Moshi
Open-source, full-duplex speech-to-speech foundation model with sub-200ms latency
LTX Video
Open-source DiT video model with synchronized audio, 4K output, and multi-keyframe control