Horch vs so-vits-svc
A side-by-side look at pricing, capabilities, pros, cons, and our editorial scores.
Horch Audio | so-vits-svc Audio | |
|---|---|---|
| Tagline | Privacy-first, on-device meeting assistant for macOS | SoftVC VITS Singing Voice Conversion — open-source pipeline for training and running singing-voice models. |
| Category | Audio | Audio |
| Pricing | Paid· One-time purchase: €49 once | Free· Free / open-source (AGPL-3.0). You provide your own compute (typically a CUDA-capable GPU) and training datasets. |
| Model | Whisper (local) for transcription; optional Ollama / MLX local LLMs for summarization | SoftVC content encoder + VITS backbone + NSF-HiFiGAN vocoder; optional ContentVec, HuBERT-Soft, Whisper-PPG, WavLM encoders and shallow-diffusion module. |
| Editorial score | — | — |
| Use cases | Confidential client meeting transcriptionAutomatic action-item extractionPer-contact relationship historyPre-meeting briefings from prior callsLocal Whisper transcription without cloud uploadFeeding meeting context into Claude or Cursor via MCPPersonal second-brain in MarkdownLegal, medical, or NDA-bound conversation notes | Singing voice conversion (AI covers)VTuber and virtual-character singing voicesCustom vocal timbre for indie music productionSpeaker mixing and timbre morphing experimentsVoice model training on curated datasetsResearch on VITS-based voice synthesisONNX export for lightweight SVC inference |
| Pros |
|
|
| Cons |
|
|
| Website | horch.app | github.com |
Pick Horch if
- ✅ Fully on-device by default — audio and transcripts never leave the Mac unless the user opts in
- ✅ One-time €49 purchase instead of a recurring per-seat SaaS bill
- ✅ Records at OS level so no bot appears in the meeting and any app (Zoom, Meet, Teams, in-person) works
- ✅ Notes stored as plain Markdown in ~/Meetings — portable, greppable, Obsidian-friendly
Pick so-vits-svc if
- ✅ Fully open source (AGPL-3.0) and runs entirely offline — no per-use fees, no data leaving your machine.
- ✅ State-of-the-art singing quality for its generation: NSF-HiFiGAN vocoder + shallow diffusion noticeably reduce breath and sibilance artifacts.
- ✅ Pluggable content encoders (ContentVec, HuBERT-Soft, Whisper-PPG, WavLM) let you trade off timbre leakage vs. pronunciation fidelity.
- ✅ Speaker mixing (static and dynamic) and clustering-based timbre control give producers real creative knobs beyond one-shot conversion.