Running Aria locally, Part 2: When things break
A home stack breaks in boring ways: the model server isn't running, the GPU is busy, a job keeps failing. How Aria shows it's offline instead of throwing errors into the chat, keeps your messages until it's back, explains failed jobs, and repairs what it safely can.
Part 1 described the stack: Aria, PostgreSQL, Qdrant and a model server you run yourself. On a home machine, at least one of those is always about to be unavailable. The model server isn’t started, the GPU is busy with a game, Docker restarted overnight, a model got unloaded.
In a chatbot, that shows up as a red error in the middle of the conversation. That’s fine for a tool. It’s jarring in something that’s meant to feel like talking to someone. So most of the work in this part went into one idea: a broken component should look like Aria being away, not like Aria being broken.
Offline, like a person
The chat header shows Aria as offline while it can’t talk: the database or the app is down, or its chat model server doesn’t answer. Hovering over it says what isn’t working. It clears on its own when things recover.
Behind that is a split between two kinds of check:
- The status check runs every probe: application, PostgreSQL, Qdrant, the memory index, the knowledge library, each model capability, the background worker. Each probe has a 5-second timeout, they run at the same time, and a probe that throws never takes the page down with it. Probe error messages are reduced to their type, so a connection string with a password can’t end up on screen.
- The online check runs only the critical probes, the ones Aria can’t talk without, and reuses the result for a few seconds. The chat header polls it every 20 seconds.
There’s also a stricter setting, on by default: OFFLINE_ON_DEGRADATION. With it, Aria also goes offline when something is wrong that would quietly harm its conversations even though it could technically still answer: the analysis or embedding model is failing, the memory index is broken, or background work on conversations keeps failing. A persona that replies fluently but has stopped forming memories is worse than one that’s honestly offline for a while, because you don’t find out until much later.
Your messages wait
Going offline would be annoying if it meant you couldn’t write. So you still can.
When a reply can’t start because the model server is down, Aria doesn’t post an error. It removes the empty reply, keeps your message, and queues the conversation. The API answers “accepted, queued” instead of failing. A background service checks the online status every 20 seconds while anything is queued, and when the model server is back, Aria answers everything you sent, together, in one reply. If the chat is open, the answer simply appears.
A failure halfway through a reply works the same way: the partial text stays, with a short note, and the reply is retried once Aria is back.
It survives restarts, too. At startup, Aria looks for conversations from the last two days that end with a message of yours that never got an answer, or with a failed reply, and queues them.
The same goes for messages Aria wanted to send on its own (the reaching-out posts cover those). If the model server is down when it wanted to write, that outreach waits instead of failing. It tries again now and then, and writes once the server answers. If several messages waited, it sends one catch-up message with whatever still makes sense, not a burst. Anything that waited more than 12 hours is dropped, because the moment has passed. Reminders you asked for are the exception: they still arrive.
One rule ties it together: model-server error text never reaches the chat. Messages and errors carry Aria’s own short notes. The raw provider error is in the logs, where it’s useful.
Failed jobs, explained
Most of Aria’s work happens in background jobs: extracting memories, indexing, summarising, analysing the relationship, reading articles. Jobs live in a PostgreSQL table, claimed with SKIP LOCKED, and failures retry with exponential back-off: 5 seconds, 10, 20, 40, and so on, capped at 15 minutes, with a little random jitter so a batch of failed jobs doesn’t retry in lockstep. After 8 attempts (about ten minutes) a job is marked failed.
A job can also say “not now”: when the thing it needs isn’t available yet, it returns to the queue for a while, the attempt isn’t counted, and the job shows as Waiting with the reason.
Early on, a failed job showed its exception and nothing else. That’s only useful if you wrote the code. The Jobs page (System → Jobs) now explains each one:
- What it was about: which persona, which conversation (with a link), which messages it still had to read, or which memory or article.
- What went wrong, in plain words: the model’s answer didn’t match the expected format (and which field), the model returned nothing, the prompt didn’t fit, the model isn’t loaded, the server couldn’t be reached, it timed out, the record no longer exists.
- What to try, with the technical error and the start of what the model actually returned underneath.
From there you can run waiting jobs now, retry failed ones (one or all) or dismiss them, and copy a report of a job to share.
A detail that turned out to matter: structured answers from models are read leniently. A number written as text (“168”, “[168]”, “#168”) is read as that number. Small local models do that more often than you’d expect, and it used to fail perfectly good jobs.
Aria knows when it’s unwell
A status page only helps if someone looks at it. So Aria checks itself.
Every 5 minutes, a self-health monitor runs the status check and keeps track of open problems: what’s not working and since when. While a problem is open, Aria’s prompt includes a short section describing it, so if you ask “is something wrong?”, it can actually tell you. It also has a tool to run a fresh check when asked.
For a few problems it knows a safe fix, and it tries it:
- Reconnect its tools when an MCP tool server dropped.
- Rebuild a missing index when a Qdrant collection disappeared, but only while Qdrant and the embedding model are healthy, and never for a collection that exists with the wrong shape.
- Retry failed background jobs from the last day, from a list of job types that are safe to retry, at most three times each, and only while the model servers and the database are healthy.
Repairs back off (1, 2, 5, 15, 30 minutes, then hourly, at most 10 attempts), and while a repairable problem is open the monitor checks every minute instead of every five. Everything else is left for me, and the System page shows the health history and every repair attempt.
A real example
While writing this series, I found a good example of why all of this exists.
Aria had been offline for a few hours (a background job had failed, and with OFFLINE_ON_DEGRADATION on, that’s enough). Two messages it wanted to send were waiting for it to come back online. Each time one of them ran, it held the other one (so they could become a single catch-up message), found Aria offline, and released the other one… immediately. The worker picked it up, it did the same thing, and the two kept starting each other with no pause at all. The attempt counter in the logs passed half a million, and the app and the database sat at high CPU doing nothing useful.
The cause was a single misplaced line: the “she’s offline, wait” path ran inside an error handler that released held jobs at once. The fix moved it out so the held jobs wait a minute, like they do when the model server is unreachable, and an integration test now runs two waiting messages against an offline Aria and checks that they back off.
It’s a nice illustration of the trade-off. Making failure quiet is good for the conversation and bad for noticing problems, so the quiet paths need their own tests, and the System page needs to show what’s waiting and why.
Seeing what it’s doing
Two more tools helped more than I expected:
- Metrics on the System page: time to first token, reply duration and outcome, context-overflow retries, provider errors by kind, retrieval time, job executions.
- A Logs page (off by default, a development aid) that streams what Aria is doing and deciding as it happens, filtered by area, level or text. Each entry opens to show its values, like the chance and the dice roll behind a decision to write first. Message content only appears there if content logging is turned on.
What I learned
- Failure should look like absence. “Offline” plus “your messages will be answered” is calmer than any error message.
- Offline beyond the obvious. Going offline when memory quietly stops working catches problems I’d otherwise discover weeks later.
- Explain errors in the user’s terms. “The model answered in a different format (the field ‘importance’)” gets fixed. A stack trace gets ignored.
- Repair only what’s safe. Reconnecting and retrying are cheap and reversible. Anything else is a decision for a person.
- Quiet paths need loud tests. The bug above did nothing visible in the chat. That’s exactly why it ran for hours.
In Part 3: moving Aria to another computer with one backup file, cloning a persona with everything it knows, and running several personas on one model server.