Jerhemy Waldon
Aria

Aria's voice, Part 1: A voice server of its own

Speaking and listening on your own GPU: Whisper for voice messages, Chatterbox, Kokoro and Orpheus behind one OpenAI-compatible speech API, a voice per persona, and a voice cloned from 5 to 30 seconds of someone speaking. Why it started on speaches and ended up as its own server.

Aria's voice, Part 1: A voice server of its own

The core of Aria is text. Everything up to Knowing itself works without a GPU of its own beyond the model server. The rest of the series is the optional extras, and the one that changes the feel the most is the first: a voice.

Hearing a reply spoken changes how it lands. It also raises the bar: a robotic voice reading a warm message is worse than no voice at all. This part is about the server that does the speaking and listening. Part 2 is about how the chat uses it.

The shape of it

Speech is two capabilities, each behind its own small interface in the application:

  • Speech to text: you press the microphone button, speak, and the transcription goes into the message box (you can edit it before sending).
  • Text to speech: each reply has a Read button, and “Read replies aloud” reads new replies automatically.

The same rule as everything else applies: the browser never talks to the speech server directly. Audio goes to Aria’s API (/api/speech/transcriptions, up to 25 MB in the usual formats), and Aria calls the speech server. Replies are made speakable on the server first: code blocks, URLs, markdown syntax and emoji are removed, because nobody wants to hear “asterisk asterisk”.

Audio is never stored or logged, in either direction.

From speaches to a server of its own

The first version used speaches, an existing OpenAI-compatible speech server, running Whisper and Kokoro on the CPU. It worked well and needed no GPU. The first voices were Piper voices per persona, then Kokoro became the default.

What it couldn’t do was clone a voice. And once I wanted Aria to have a voice of its own, not one picked from a list, that was the missing piece.

So Aria got its own voice server: a small FastAPI service on the GPU (the same approach as the image server) that runs interchangeable engines, one loaded at a time:

Engine What it’s good at
Chatterbox (and Chatterbox Turbo) Natural, expressive speech, and cloning from a short sample, with an emotion intensity setting. MIT licensed.
Kokoro Small, fast and clear. The default voice.
Orpheus Another expressive engine, for comparison.

The trick that made it easy to adopt: the voice server speaks the OpenAI speech API (/audio/speech) and answers the same voice-registry calls speaches did. Aria’s speech client didn’t need a rewrite, just a second address and a cloning call. Model names from other servers (speaches’ Kokoro, Piper) map to the voice server’s engine of the same kind, so voices saved before the switch kept sounding the same.

Then speech to text moved over too: Whisper, through faster-whisper, on the GPU, behind the OpenAI transcription endpoint. At that point speaches had nothing left to do, and it was removed. One speech server, one Compose profile (voice).

A voice per persona

Each persona has its own voice. The voice picker on its profile lists what the voice server offers (cached for ten minutes, falling back to the default if the server is down), with a preview, so you can hear a voice before choosing it.

Cloning from a sample

The feature that started all this: “Clone her voice from a sample” on a persona’s profile. Upload 5 to 30 seconds of someone speaking, and the persona speaks in that voice from then on, through Chatterbox.

The sample is stored on the voice server under the persona’s name, and spoken by the cloning engine. Clean audio of a single speaker works best. And the obvious rule applies: clone voices you have the right to use.

Choosing a persona's voice, or cloning one from a sample

Why it stays a separate server

When I later combined the image server and the talking-picture server into one avatar server (the looks and talking picture posts cover those), the natural question was whether the voice should join too. One GPU service is easier to manage than three.

It can’t, for a boring and very real reason: dependencies. Chatterbox pins PyTorch 2.6 and NumPy 1.x. The image models and the talking picture run on PyTorch 2.8 and NumPy 2. They can’t share a Python environment. So there are three images, each named for what it is: aria-server (the app), aria-voice and aria-avatar.

Keeping it apart has an upside too: the voice server can run on a different machine from the avatar server, each behind its own API key.

Sharing one GPU

On my machine, the chat model, the voice server and the avatar server all share one 24 GB GPU. I measured it with everything loaded at once: 22 of 23 GB in use, every request succeeded, and the talking-picture renders were slower but fine. Each server unloads its model after a period of idle, so in practice they rarely all hold memory at the same time. No scheduler between the servers was needed.

Pronunciations

One problem showed up immediately: names. My name, Jerhemy, is said like “Jeremy”, but a voice engine reads it the way it’s spelled.

So Aria has pronunciations: a per-user list mapping a written word to how it’s said. Synthesis replaces whole words (in any case) after making the text speakable, so only the voice changes. The text keeps the real spelling, and every reply’s prompt includes a short spelling note so the model keeps writing “Jerhemy”, not “Jeremy”.

You manage them in Settings, and Aria can save one itself when you tell it how something is said (“it’s pronounced SHIV-awn”), with a tool.

The voice server’s own page

The voice server has its own settings page, a small React app served by the server itself (the same stack as Aria’s chat app):

  • Its voice: pick and load the engine, play, clone or delete voices, and try voices side by side with the tries kept to compare.
  • Listening: choose the Whisper model and try a transcription.

The server also gained endpoints to load or unload an engine, and timing headers on spoken audio, which helped tune the next part.

The voice server's own page: engines, voices and tries

What I learned

  • Speak the standard API. Making the voice server look like the OpenAI speech API meant Aria’s client barely changed, and the old voices kept working.
  • Dependencies decide architecture. The cleanest-looking design (one GPU service) was blocked by two versions of NumPy. Separate images were the honest answer.
  • Measure shared GPUs before building a scheduler. Idle unloading plus fallback modes covered it. A scheduler would have been complexity for a problem that didn’t happen.
  • Names are the first thing a voice gets wrong. Pronunciations that change only the audio fixed it without touching the text.

In Part 2: when Aria’s voice leads. Speaking each sentence as soon as it’s written, captions that follow the voice word by word, and waiting for a picture before speaking.