Skip to main content
📖 The AI Tool Bible

Dia vs Horch

A side-by-side look at pricing, capabilities, pros, cons, and our editorial scores.

 Dia logo
Dia
Audio
Horch logo
Horch
Audio
TaglineOpen-weights 1.6B text-to-dialogue model that generates ultra-realistic multi-speaker conversations in one pass.Privacy-first, on-device meeting assistant for macOS
CategoryAudioAudio
PricingFree· Free, open weights (Apache 2.0); hosted larger version waitlistedPaid· One-time purchase: €49 once
ModelDia-1.6BWhisper (local) for transcription; optional Ollama / MLX local LLMs for summarization
Editorial score7.3 / 10—
Use cases
dialogue-generationvoice-cloningpodcast-prototypinggame-voice-actingtext-to-speech
Confidential client meeting transcriptionAutomatic action-item extractionPer-contact relationship historyPre-meeting briefings from prior callsLocal Whisper transcription without cloud uploadFeeding meeting context into Claude or Cursor via MCPPersonal second-brain in MarkdownLegal, medical, or NDA-bound conversation notes
Pros
  • Open weights under Apache 2.0 with first-party Transformers support
  • Multi-speaker [S1]/[S2] dialogue and nonverbal tags in a single pass
  • Zero-shot voice cloning from a short audio prompt plus transcript
  • Runs ~2x realtime on a single RTX 4090 at ~4.4GB VRAM
  • Free Hugging Face ZeroGPU Space to try without local GPU
  • Fully on-device by default — audio and transcripts never leave the Mac unless the user opts in
  • One-time €49 purchase instead of a recurring per-seat SaaS bill
  • Records at OS level so no bot appears in the meeting and any app (Zoom, Meet, Teams, in-person) works
  • Notes stored as plain Markdown in ~/Meetings — portable, greppable, Obsidian-friendly
  • Auto-generates action items, topic summaries, and per-person profiles from spoken content
  • MCP server exposes meeting history to Claude, Cursor, and other agent clients
  • Local Whisper transcription plus optional Ollama/MLX means you pick the summarizer
Cons
  • English only; no built-in multilingual support
  • Voices drift between runs unless you fix a seed or supply a prompt
  • GPU required; CPU inference not yet supported
  • Tiny team (1.5 engineers); slower issue turnaround than commercial TTS
  • macOS only — no Windows, Linux, iOS, or web client
  • Local Whisper and summarization need a reasonably modern Apple Silicon Mac to feel fast
  • No cloud sync or team workspace, so sharing across a team requires bring-your-own storage
  • Small independent product without the integrations catalog of Fireflies, Otter, or Fathom
  • OS-level capture depends on macOS screen/audio permissions, which some corporate MDM setups block
Websitegithub.comhorch.app
Pick Dia if
  • ✅ Open weights under Apache 2.0 with first-party Transformers support
  • ✅ Multi-speaker [S1]/[S2] dialogue and nonverbal tags in a single pass
  • ✅ Zero-shot voice cloning from a short audio prompt plus transcript
  • ✅ Runs ~2x realtime on a single RTX 4090 at ~4.4GB VRAM
Pick Horch if
  • ✅ Fully on-device by default — audio and transcripts never leave the Mac unless the user opts in
  • ✅ One-time €49 purchase instead of a recurring per-seat SaaS bill
  • ✅ Records at OS level so no bot appears in the meeting and any app (Zoom, Meet, Teams, in-person) works
  • ✅ Notes stored as plain Markdown in ~/Meetings — portable, greppable, Obsidian-friendly