

Harmonai
Open-source generative audio lab from Stability AI building diffusion models for music production.
In short
Harmonai is Stability AI's open-source lab for audio diffusion models like Dance Diffusion. It targets ML engineers and producers wanting to run or fine-tune local audio generation tools.
Pick Harmonai if you're a producer or ML engineer who wants to run open-source audio diffusion models locally and fine-tune them on your own samples.
Skip it if you want a polished, hosted text-to-music app with a UI and account — use Stable Audio or Suno instead.
Harmonai is a Stability AI research lab that develops and releases open-source generative audio models aimed at music producers and sound designers. The group is best known for shipping Dance Diffusion, an early diffusion-based music generator, and for contributing to the Stable Audio family of text-to-audio models. Outputs include raw waveform generation, custom sample libraries, and experimental tools for building infinite, royalty-free sound material.
It's not a polished SaaS product with a billing page — it's a lab. The website is a portal pointing to a GitHub org and a Discord community where models, training code, and Colab notebooks are released. That makes Harmonai a fit for technically inclined musicians, ML researchers, and audio tool builders who want to run or fine-tune models locally, rather than for consumers looking for a one-click music generator.
If you want a hosted product layer on top of similar tech, Stability AI's Stable Audio service is the commercial sibling. Harmonai itself is the upstream research and open-weights side of that pipeline.
Harmonai is the research wellspring behind a lot of modern open audio generation, and the open weights matter. Just don't expect a product — expect a GitHub org and a Discord. For the right user that's the appeal, not a flaw.
— The AI Tool Bible editorial team
Pros
- ✅ Genuinely open-source weights and code under a real research lab
- ✅ Backed by Stability AI with serious audio-diffusion expertise
- ✅ Useful for fine-tuning custom sample libraries and unique sound design
- ✅ Active Discord and GitHub community around the models
Cons
- ⚠️ Landing page is sparse; you need to dig into GitHub to find tools
- ⚠️ No hosted UI or one-click product for non-technical users
- ⚠️ Release cadence is research-paced, not product-paced
Use cases
Frequently asked
- Is Harmonai free to use?
- Yes, Harmonai offers free open-source models and code. There is no hosted product or billing page on the site; it functions as a research lab releasing weights and training code for local use.
- Who is Harmonai best suited for?
- It is ideal for technically inclined musicians, ML researchers, and audio tool builders who want to run or fine-tune open-source audio diffusion models locally. It is not designed for consumers seeking a one-click, polished SaaS experience.
- What models does Harmonai provide?
- The lab is known for shipping Dance Diffusion, an early diffusion-based music generator, and contributing to the Stable Audio family of text-to-audio models. These tools support raw waveform generation and experimental sound material creation.
- Is there a hosted version of Harmonai?
- No, Harmonai is not a polished SaaS product. It is a portal pointing to a GitHub org and Discord community. For a hosted commercial product using similar tech, Stability AI offers the Stable Audio service.
- What are the main use cases for Harmonai?
- Primary use cases include music generation, sound design, sample library creation, audio research, and model fine-tuning. It allows users to build infinite, royalty-free sound material using open-weights models and Colab notebooks.
Explore related
Compare with similar tools
All in Audio →ElevenLabs
FeaturedThe gold standard for AI voice cloning and TTS.
Suno
FeaturedText-to-song AI — full vocal tracks from a prompt.
Udio
Suno's main rival for AI-generated full songs.
AssemblyAI
Speech-to-text API with diarisation, summarisation, and topic detection.
Chorus by ZoomInfo
Enterprise conversation intelligence bundled with ZoomInfo's B2B data graph
Whisper
OpenAI's open-source speech-to-text — the de-facto baseline.