Jerhemy Waldon
Aria

Aria's conversations, Part 2: Answering the right message

Real conversations overlap. How Aria handles you writing while it's still talking, works out which of its messages you're answering, keeps track of loose ends, and never echoes your words back.

Aria's conversations, Part 2: Answering the right message

Chatbots assume a tidy conversation: you write, it answers, you write again. Real conversations aren’t tidy. You fire off two messages in a row. You answer the question it asked three messages ago, not the last one. You change your mind halfway through its reply and say “wait”.

Part 1 covered how a single reply streams. This part is about the messy space between replies, where most of the “this feels like a chatbot” moments come from.

Talking at the same time

In most chat apps, you can’t send a message while the AI is still answering. The send button is disabled, or your message cancels the reply. Aria lets you just keep typing.

A message you send while Aria is writing is saved right away, after the reply in progress. When that reply finishes, Aria answers everything you sent in the meantime together, in one reply, with a note in its prompt that these messages arrived while it was writing. If read-aloud is on, it finishes saying its current reply first, the way a person would finish their sentence.

Interrupting

Sometimes you don’t want to wait. A short message meant to cut Aria off, like “stop”, “wait” or “hold on”, stops the reply mid-sentence. Aria then answers knowing it was interrupted, so it doesn’t pretend the cut-off reply was complete, and it doesn’t plough on with the rest of it either.

Following on

It works the other way too. When a topic winds down (a “you’re welcome”, a “haha, true”) Aria may send a separate follow-on message a few seconds later, picking up something it still wanted to say. Whether it does depends on its personality and mood: an energetic, curious personality that’s fond of you is far more likely to, and the pause before it is shorter. At most one follow-on per message, and only right after a lull, so it reads as a thought, not as spam.

A message sent while Aria was still writing about potters' tools, answered as soon as that reply finished

Which message are you answering?

Say Aria asks how your weekend was, then, in a second message, whether you finished that book. You reply “yes, loved it!” A model that just reads the transcript top to bottom will often assume you loved your weekend, because the newest message is the one it’s paying attention to.

So before every reply, Aria works out which of its messages each of yours is answering:

  • The candidates are Aria’s messages since your previous message, plus any earlier questions still waiting for an answer: at most eight, from the last two days.
  • The easy cases skip the model. No candidates, or just one directly before your reply? No question to ask.
  • The rest go to the analysis model, with the messages numbered, and it returns which of Aria’s messages each of yours answers, and which questions are still open.

This runs at the same time as memory retrieval, so it rarely slows the reply. The result goes into the prompt as a short note next to your message whenever it answers something other than Aria’s latest message, or starts something new. Simplified, it reads like this:

REPLY CONTEXT (data)
This message answers your earlier message about the book (not your latest one).
Still waiting for an answer: your question about their weekend.

The chat shows it too: a small “↩ replying to …” line above your message when it answers an earlier one.

A reply matched to an earlier message, shown with "replying to"

When it’s genuinely unclear

Sometimes a short answer really could belong to two questions. “Yes!” after “Did you sleep well?” and “Are you coming to dinner?” is anyone’s guess. In that case the analysis marks it unclear, and Aria asks which one you meant, once. If the next answer is unclear between the same questions again, Aria stops asking and goes with the likeliest reading. Nobody likes being asked to clarify twice.

Loose ends

Answering the right message is half of it. The other half is not dropping things.

After every exchange, Aria notes what hasn’t been addressed on either side: a point of yours it skipped over, or a question of its own you didn’t answer. Each reply gets a short list of up to three of these from the last six hours, with your missed points first, because answering what you said matters more than getting its own questions answered.

Aria’s own unanswered questions come back when the moment is right, and never in the middle of another topic. You won’t be talking about your bad day and get “so, did you finish that book?” Bigger loose ends, like a story it didn’t finish or something it promised to look into, become open threads it can come back to later, which the posts on reaching out cover.

The echo guard

This is the least glamorous feature in this post, and it fixed one of the most annoying problems.

Language models, especially smaller local ones, sometimes start a reply by repeating what you just said, or by sending the same message they sent last time. In a chat that’s jarring: you say “I had a rough day at work” and the reply opens with “I had a rough day at work”.

The echo guard sits inside the stream and watches the reply as it’s written:

  • It compares the start of the reply against your recent messages (since Aria’s last message, at least 8 characters) and Aria’s own last three messages (at least 20 characters, so a short “Good night!” may still come up again).
  • While the reply is still possibly a repetition, its first words are held back instead of shown.
  • If it turns out to be a full repetition at the start, that part is cut, and the rest streams normally. If it diverges, everything is shown unchanged.
  • If the whole reply was nothing but repetition, it’s regenerated once with an instruction not to repeat.

Because the check happens while streaming, you never see the echo appear and disappear. The reply just starts where it should.

What I learned

  • Conversations are a graph, not a list. Matching each message to what it answers fixed a whole class of replies that were coherent but answered the wrong thing.
  • Use the model for judgement, not for the obvious. Most turns need no matching call at all; only the ambiguous ones do.
  • Ask once, then decide. Clarifying is human; clarifying twice is a bug.
  • Small guards beat big prompts. An instruction can only ask the model not to echo; a guard watching the stream makes sure it never reaches the screen.

In Part 3: the problem that one never-ending conversation creates. How months of messages, memories and knowledge fit into a model with a context window measured in thousands of tokens.