VoiceStudio — open-source ElevenLabs replacement (Dr. Alvaro Cintas)

Author: Dr. Alvaro Cintas (LinkedIn)

ElevenLabs just got a full open-source replacement that runs entirely on your own hardware.

It’s called VoiceStudio, and it does the entire voice pipeline locally: cloning, voice design, video dubbing, dictation, transcription, and audiobook creation.

Under the hood it routes across 16 TTS engines and 11 ASR engines, with:

  • Voice cloning — zero-shot from a 5–15 second reference clip, no training required
  • Video dubbing — transcribes, translates, keeps the original speaker’s voice, re-exports the video
  • Dictation — a system-wide shortcut that transcribes live and can clean the text with a local LLM
  • Audiobooks — multi-voice scripts with EPUB/PDF import, rendered chapter by chapter into .m4b

It also ships an OpenAI-compatible audio API.

Notable comments

What stands out to me is less the local hosting and more that VoiceStudio is an orchestration layer over 16 TTS and 11 ASR engines, not one model doing everything. That routing decision matters more long term than which single engine wins this month, since engines keep leapfrogging each other while an interface that picks the right one per job doesn’t need to change.

Just a heads up, it’s good but not the best out there, and a bit of a resource hog.

I was going to do this. Glad someone actually did.

This gives startups a much better path to scale. A SaaS API often makes more financial sense early on, another point to use ElevenLabs.

This looks cool.

I wonder how soon our “digital representatives” will start talking to each other, realize they’re both agents, and switch to bytecode :)