Aria's talking picture, Part 1: Its picture speaks
Animating a still portrait to speak each reply, on a single home GPU, in about real time. Why Ditto, how speedups took it from 18 to about 24 frames per second, sentence-by-sentence versus whole-reply videos, a draggable window, and why it always uses the same face.
Aria has a voice and a face. The obvious next step was to put them together: Aria’s portrait, speaking its replies.
This is the most eye-catching feature in the project, and also the most optional. It needs a GPU, the voice server and the avatar server. Without them, Aria works exactly the same. With them, replies play as a short video of its picture talking.
Choosing a model
The requirements were specific:
- One still picture in, not a video of a real person or a trained avatar per face. A persona has a portrait and nothing else.
- About real time on one home GPU, shared with the chat model and the voice. If a sentence takes ten seconds to render, the voice has long finished.
- Open weights and a usable license, running locally like everything else.
- Practical to deploy: no per-GPU engine builds.
I looked at the well-known options. Wav2Lip moves the mouth but leaves the rest of the face frozen, which looks uncanny on a portrait. SadTalker animates the whole head but isn’t built for real-time use. Ditto (from Ant Group, Apache 2.0) makes a still portrait speak in about real time, animates the head and face as well as the mouth, and runs from plain PyTorch weights without building TensorRT engines for each GPU.
Making it fast enough
Out of the box, Ditto’s default of 50 diffusion steps for the motion ran at about 18 frames per second on an RTX 4090. Video plays at 25. That means a sentence takes longer to make than to say, and the gap grows with every sentence.
Two changes got it to about 24 frames per second:
- Fewer steps. The motion is generated with 10 diffusion steps instead of 50 (
TALKING_SAMPLING_STEPS). For a talking portrait, the difference is hard to see. - Keeping work on the GPU. Ditto moved some intermediate features (the warped feature volume, and the portrait’s own features) between stages in a way that cost time. A small set of speedups, applied when the model loads without changing Ditto’s own files, keeps them on the GPU between stages.
At that speed, a sentence takes about as long to make as to say, which is exactly what sentence-by-sentence playback needs.
How a reply becomes a video
Aria does the coordinating. The voice server and the avatar server never talk to each other:
- Aria synthesises a segment of the reply in the persona’s voice (any engine, cloned or not).
- It sends the picture and the audio to the avatar server’s talking endpoint.
- The avatar server returns an MP4 of the picture speaking that audio.
The chat asks for one video per segment and prepares the next while one plays. If a video can’t be made, the rest of the reading switches to plain audio. The reply is still spoken, just without the picture.
Videos are kept in memory for a day (at most 64 MB, about 200 sentences), keyed by the user, the persona, the voice, the picture and the exact text as spoken. Reading a reply again replays at once without a new synthesis or render. Like all audio in Aria, they’re never written to disk.
Sentence by sentence, or the whole reply
There are two ways to cut a reply into videos, chosen in Settings:
- Sentence by sentence. The same segments as voice-led reading: speaking starts soon after the first sentence is written. The catch: each sentence is its own take, so the head returns to its starting pose between sentences.
- The whole reply at once. The chat waits for the finished reply and asks for as few videos as possible (up to 1,500 characters each, about 100 seconds of speech), so the head moves continuously. It takes longer to start, and the window shows that it’s getting ready.
Code blocks still split a reply in both modes, because code is shown, not spoken.
Always the same face
The first version used the expression picture of each reply, so a happy reply was spoken by the happy portrait. In practice that was worse. Expression pictures are separate drawings, so the face changed slightly from reply to reply, and a talking face that keeps changing shape is unsettling.
So the talking picture now always uses the persona’s base portrait, or a picture uploaded specially for talking on its profile. The emotion comes from the voice and from motion settings (Part 2), not from swapping pictures. A bonus: a cached video now fits any later reply with the same text.
Where it plays
The talking picture went through three layouts in quick succession:
- A column on the right of the chat. It took a third of the chat’s width.
- A toggle in the chat header, a camera button next to “Read replies aloud”, remembered per browser, instead of a setting per persona. That part stayed.
- A movable window. The video now plays in a small window over the chat that you drag by its title bar (or with the arrow keys) and resize from its corner. It stays square, at least 120 pixels, always fully on the page, and it remembers where it was. Between replies it shows the same picture, still. Closing it turns the talking picture off.
The window is plain pointer events in the app. With only one window, a window-management library would have added more than it saved.
One GPU for everything
On my machine, the chat model, the voice server, the pictures and the talking picture all share one 24 GB GPU. With everything loaded at once, memory sat at 22 of 23 GB, every request succeeded, and renders took 5 to 9 seconds instead of about 5 for 4 seconds of speech.
Pictures and the talking picture later moved into one avatar server, which helps here: one lock means a picture and a talking video never render at the same time, and when the GPU is short of memory, loading one part unloads the other. The voice server stays separate for dependency reasons (Voice, Part 1).
One more fix from measuring: render times now exclude the time spent waiting for the GPU lock, so the timings show how long rendering actually takes.
What’s next
Ditto can also start from a short video instead of a still. The parts the audio doesn’t drive (breathing, small head movements, blinks, light) would then come from real footage instead of being invented from one frame. The mouth would still follow the speech. That’s planned, not built.
What I learned
- Real time is the requirement that picks the model. Quality differences between talking-head models mattered less than whether a sentence could be made as fast as it’s said.
- Profile before you buy a bigger GPU. Fewer diffusion steps and keeping tensors on the device gained about a third in speed.
- Consistency beats expressiveness. One steady face looked more alive than a face that changed with every reply.
- Try layouts in the real app. A column looked right in a sketch and felt wrong in use. A movable window was the third try, and the one that stuck.
In Part 2: tuning how it moves, on the avatar server’s own page, with a test bench to compare settings before saving them.