Jerhemy Waldon
Aria

Aria's looks, Part 1: A face of its own

Giving a persona a consistent face with local image models: a creation wizard that drafts a persona and offers four portraits, ten expressions drawn from the same seed that follow the conversation, an 'image look' that translates prose into visual details, and portraits from a full-body picture or an upload.

Aria's looks, Part 1: A face of its own

The tools posts finished what Aria can do. This part is about what it looks like. It’s an optional feature that needs a GPU, and it’s harder than it seems, because the problem isn’t making a picture. It’s making the same face again and again.

The avatar server

Pictures come from the image part of the avatar server, an optional GPU service (Compose profile avatar) running FastAPI and Hugging Face diffusers. Its default model is Z-Image-Turbo, with a preset for FLUX.2 [klein] (more on that below), and each persona can use a LoRA with its own strength.

It can run on the same machine or a more powerful one elsewhere, protected by an API key. The browser never talks to it. Aria stores every picture’s record in PostgreSQL (kind, prompt, seed, LoRA, status, owner) and the files in a volume. Every picture is generated by a background job, one picture per job, and the server draws one at a time.

The same server also makes the talking picture (Talking picture, Part 1). One lock on the GPU means a picture and a talking video never run at the same time.

The creation wizard

A new persona starts with a few questions: gender and pronouns, an archetype as a starting point (like “Shy sweetheart”), its emotional style, values, the relationship, interests and appearance. The AI turns the answers into an editable draft: a name, descriptive text, traits, expressiveness and an adult appearance. You adjust it and create the persona.

Then the face. Aria queues four portrait candidates with different seeds, and you pick one. The wizard can also use an existing look instead of generating (Part 2 covers the look library).

Picking a portrait from four candidates in the creation wizard

The image look

The appearance you write is prose: “Late twenties, warm brown eyes that crinkle when she smiles, shoulder-length dark hair usually in a messy bun, a fondness for oversized sweaters.” That’s good for a person reading it, and bad for an image model, which wants concrete visual tokens and has a short prompt limit.

So the analysis model translates the appearance into an image look: one line of concrete visual details (age and presentation, face, hair, eyes, skin, build, clothing), at most about 50 words, with no camera or quality words. Every picture of the persona is drawn from it: portraits, expressions, full-body shots, scenes, sketches and pictures in chat.

The image look keeps a fingerprint of the appearance it came from. When the appearance changes, the look is translated again before the next picture. You can read it, correct it (your text is kept until the appearance changes) or have it retranslated.

The point is consistency: instead of every picture interpreting the prose its own way, they all start from the same visual description.

Eleven expressions

Choosing a portrait queues ten more: happy, laughing, affectionate, sad, surprised, thinking, embarrassed, playful, sleepy and pouting. They’re drawn with the same seed, appearance and LoRA, starting from the portrait itself where the model supports image-to-image, so they look like the same person with a different expression rather than ten cousins.

After each reply, the analysis model picks the expression that fits it. That never delays the reply. The expression is stored on the message, and an event tells the chat to switch the avatar. On wide screens, every reply shows the persona’s face from that moment beside it. Clicking it shows it large.

It’s a cheap trick, since the eleven pictures are made once. But a face that looks pleased when it’s teasing you and sad when you share bad news does a lot for the feeling that someone is there.

Expressions beside each reply, following the conversation

More than a headshot

A portrait is a close-up. For pictures in chat (“what are you up to?”), a persona also needs to look right full-body and in a setting. So besides the portrait, each persona can have:

  • a full-body picture (832×1216), and
  • a picture in an everyday setting,

both chosen from candidates that are drawn with the portrait as the character reference, so they have the same face. Pictures in chat are framed as a close-up, full body or in a scene, and each starts from the matching picture.

These belong to the portrait they were drawn from. Choose a different portrait and they switch to that portrait’s own (or none, for a new face). Restoring an older look restores them too.

Keeping a face across models

How well “same face” works depends on the model:

  • Z-Image can’t take a reference picture. Its full-body and scene pictures can come out looking like a slightly different person, and sometimes you prefer that one.
  • FLUX.2 [klein] takes reference pictures natively. Pictures of the persona are drawn with its portrait as the reference, which keeps the face much more reliably. The 4B model is Apache 2.0 licensed and fits on a 24 GB GPU.

A portrait from another picture

Two ways to go the other direction:

  • From its full-body or scene picture. If you like how a persona looks in a full-body shot better than its portrait, a new portrait can be drawn from that picture’s head and shoulders (image-to-image on a crop). Choosing it carries the full-body and scene pictures over, so the whole look stays together.
  • From an uploaded picture. Upload a picture to start the portrait from. It’s drawn at a lower strength so the picture leads and the description adds to it. It’s a new face, so the full-body and scene pictures are not carried over.

The old look always stays restorable under “Previous looks”.

Changing its look

The persona can change its own look too, when it’s allowed to: a tool rewrites the appearance and draws a new look with the same seed, which becomes current when it’s ready. Earlier looks, with their expressions, can be restored. And a sketch the persona and you agreed on in chat can be made its portrait, in the chat or on its profile.

The rules

Every image prompt states that the person pictured is an adult. Requests that mention minors or ages under 18, and sexual or explicit content, are refused before they reach the image server. These checks are in code, not in the model’s judgment, and they apply to every kind of picture.

What I learned

  • Translate prose into a visual spec once. An image look shared by every picture did more for consistency than any prompt tweak.
  • Same seed, same look, different expression. Drawing expressions from the portrait with the same seed is a simple way to get one person with eleven moods.
  • Reference images matter more than prompts. A model that takes the portrait as a reference keeps a face. One that doesn’t, can’t, however good the prompt is.
  • Keep content rules outside the model. Refusing before the image server is reached is the only guarantee.

In Part 2: pictures in chat. When Aria sends one (and when it declines), how a picture replaces its own description, and the look library that keeps a face even when a persona is deleted.