Version 2.0 · Coming Soon

Audiobook Studio 2.0

Casting for every character, plug-in voice engines, and a sturdier queue.

  • Runs on your machine
  • No subscription
  • No per-word bill

See what's new

Poke around the 2.0 interface.

It runs on sample data, so click anything: library, voices, engines, settings. Nothing is uploaded or saved.

Audiobook Studio 2.0 · interactive demo
Hear the voice

Same words, two engines.

Both clips use the Studio Narrator voice and read exactly the same text. One is made on your own computer, the other with a cloud engine.

Local · default

XTTS

Runs on your computer. Private and free.

Cloud · optional

Voxtral

Turn it on with your own API key. It sends the text and any reference audio to Mistral.

Engines are plugins, so more can be added.

Read the transcript

Welcome to Audiobook Studio. This sample uses the same voice profile with two different synthesis engines so you can hear how the delivery changes while the script stays exactly the same. Audiobook Studio is centered on a local-first, private workflow. XTTS is the fully local engine. Voxtral is the optional cloud engine, which means Voxtral generation sends synthesis text and selected reference audio to Mistral.

The goal here is not to prove that one engine wins every sentence. It is to show tone, pacing, clarity, and how each engine handles a calm, explanatory read over multiple lines of connected narration. If you are listening closely, pay attention to the transitions between phrases, the weight of emphasized words, and the natural rise and fall at the ends of sentences.

Coming in 2.0

What's new.

We'll announce it on this page when it's ready.

Browse the full 2.0 docs. Staying on 1.x for now? The 1.x guide is here.

A voice for every character

Cast each character with their own voice, and mix engines inside a single chapter.

Casting suggestions

Accept or change suggested matches, so filling out a cast goes faster.

Voice engine plugins

Install a new engine straight from a GitHub repository, or build your own with the plugin SDK.

Voice library and bundles

Portable voice bundles you can import from and export to Hugging Face.

A sturdier queue

The text-to-speech server runs on its own, so one engine crashing doesn't take down the app, and the queue picks up where it left off after a restart.

Gateway API

Use Studio as a local text-to-speech gateway for your own tools, with built-in OpenAPI docs.

2.0 is built in the open. Curious how it works? View on GitHub

The fine print

Good to know.

  • Free and open source. Version 2.0 will be released under the Apache-2.0 license.
  • The XTTS voice model is non-commercial. Audio made with it can't be sold.
  • Version 2.0 will include Studio Narrator. An AI voice created from the voice of Steven L. Dunn. For commercial use of the voice itself, get in touch.
  • The demo runs on sample data. It's a preview of the 2.0 interface, not the finished release.