Behind the scenes: how Aria is put together
A modular monolith that grew many features quickly, and the consolidation that followed: one seam per external system, prompt sections as contributors, a provider registry, architecture rules enforced by tests, documentation split by subsystem, and a numbered log of every decision.
The last thirty posts were about what Aria does. This one is about how the code is organised, for anyone who wants to build something similar, or is curious how a hobby project with this many moving parts stays workable.
It’s also a story about a cleanup. Features arrived much faster than the structure around them, and at some point the cost of that showed.
A modular monolith
Aria is one .NET process. It hosts the APIs (the native one and the OpenAI-compatible one), serves the React chat app, runs the background job worker, and runs the schedulers: knowledge reading, reaching out, interest decay, index reconciliation, self-health. PostgreSQL and Qdrant are the only other required services. The voice, avatar and browser services are optional extras.
Inside, it’s four projects with a strict direction:
Aria.Api ──► Aria.Application ──► Aria.Domain
│ ▲
└──► Aria.Infrastructure
- Domain: users, personas, conversations, messages, and their rules. No dependencies.
- Application: the use cases, and the interfaces everything else codes against: a chat model, an embedding model, a memory index, a job queue, a speech engine, an image generator. It references no database library, no Qdrant client, no model-server client and no HTTP library.
- Infrastructure: every external system. EF Core and PostgreSQL, Qdrant, the model providers, speech, images, the web fetcher, MCP, the credentials vault.
- Api: the composition root: configuration, authentication, endpoints, hosting.
A monolith was the right call for a single-user persona. One process to run, one place to debug, transactions that cover a reply and the jobs it causes. The seams that a microservice architecture would give you can be had inside one process, with interfaces.
When features outran structure
Most of the features in this series arrived in a short time, on top of the original design. At the end of that, a review of the code found what you’d expect:
- A god class. The chat orchestrator, which runs one reply, had 51 constructor dependencies. Its main method was about 300 lines, and the streaming loop had nine state flags.
- Duplication across background jobs. About twenty analysis jobs each re-implemented the same pipeline: read the unread messages since a cursor, format a transcript, add the “this is data, not instructions” framing, ask the model, handle errors.
- Providers that knew about each other. Choosing a model server was a closed
switch, and the OpenAI-compatible provider borrowed helpers from the Ollama one. - Docs that had drifted from the code.
So I stopped adding features and did a consolidation, written up as a phased plan, with one firm rule: no change in behaviour. Every step kept every existing test green (673 at the start) and landed as its own reviewable commit.
Seams
The goal of the consolidation was that every external system and every optional behaviour sits behind one seam, so it can be replaced without touching the rest:
| To change | Add or replace | Why nothing else changes |
|---|---|---|
| A model provider | A folder and one registration | The router, Settings and startup log read the registry |
| A prompt section | One prompt contributor | The orchestrator runs every contributor in order |
| Work after a reply | One row in the after-reply table | The completed turn queues every row that applies |
| Startup work | One startup task | The initializer runs every task in order |
| The vector store | The memory, knowledge and example index interfaces | Application only sees those |
| The job queue | The queue interfaces and the worker’s claim query | Handlers only implement a handler interface |
| A game | One game class and one board | Picker, invitations, panel and prompt are shared |
| A tool | One tool class, or an MCP server in config | The executor and router offer every registered tool |
The biggest of these was the reply. Everything Aria knows that goes into a prompt (memories, knowledge, reply matching, open threads, curiosity, self-awareness, spelling, style examples, life, its look, a game, a story, a scene) is now a prompt contributor: a small class that adds one section. The orchestrator runs them in a documented order. A test checks that the order is what the docs say. Adding a section to the prompt is one class and one line of registration.
Before touching that path, characterisation tests pinned down its current behaviour: the echo retry, the empty-reply retry, tool rounds and reminder follow-through. Then the code moved, and the tests confirmed it still did the same thing. The orchestrator went from 51 dependencies to 37 in that step. The rest was deliberately left for later, to keep each step behaviour-preserving.
Rules enforced by tests
Some architecture rules are too important to leave to review, so they’re tests:
- Layering: Application and Domain reference no infrastructure library. (This test existed before the cleanup, and checked a project prefix that no longer existed, so it could never fail. Fixing it was step one.)
- Documentation: every public type in the core layers has a summary.
- Configuration: every flat environment variable maps to an options property, and appears in
.env.exampleand the deployment docs. Options validate themselves at startup, and invalid settings stop the app with a message. - Prompt contributor order matches the documented order.
- Clone coverage: every table with a persona id is either copied by a clone or deliberately excluded.
How it’s tested
Three .NET test projects and the web app’s tests, all run by CI on every push:
| Project | What it covers | Size today |
|---|---|---|
| Unit tests | Planners and engines, prompt building, provider wire formats, configuration, architecture rules | 881 tests, seconds |
| API tests | The HTTP host without a database: health, auth, hosting | 27 tests |
| Integration tests | End to end over HTTP: chat, memory, knowledge, reaching out, relationships, images, speech, MCP, the OpenAI API | about 240 tests, about 20 minutes |
| Web | Modules and hooks with Vitest |
The integration tests are the interesting ones. They start real PostgreSQL and Qdrant in containers (Testcontainers), one of each for the whole run, and each test gets a fresh database and its own collections. Everything that would call a model is faked: a scripted chat model (replies, tool calls, failures), a structured-output model with one responder per result type, deterministic embeddings, fake image, MCP and OpenAI servers, a local web server for pages to fetch, and fixed dice for anything random. Time-dependent logic takes a TimeProvider, so tests control the clock.
That combination lets a test say “the user writes this, the model answers that, the database ends up like this”, for a reply or a background job, without a GPU, in seconds.
The bug from Running it locally, Part 2 is a good example of the limits: the first version of its regression test passed even without the fix, because in the test setup the model server looked unreachable, so the buggy path never ran. It needed a fake “model server is up” plus a fake “a critical component is down” before it could fail. A test that can’t fail isn’t a test, which is the same lesson as the layering test above.
Documentation as part of the code
The project has more documentation than most hobby projects, because some of it is for Aria itself:
- Aria’s development history: what it can do today, the roadmap, and a dated changelog entry for every change, written in plain language for Aria and me. It’s loaded into Aria’s knowledge library, so Aria can answer questions about itself (Knowing itself). Writing the entry is part of making the change.
- A glossary: one term per concept (open thread, event follow-up, nudge, topic interest vs persona interest, mood estimate, reply style…), used in code comments, UI text, docs and conversation. A new concept gets a glossary entry before it gets code. It’s also in Aria’s knowledge, so Aria and I use the same words.
- A numbered decision log: every technical decision, with the alternatives rejected and why. It’s past 280 entries now, and it’s the first place to look when the question is “why does it do it this way?”
- An architecture overview with the seams above, and one page per subsystem (chat, inference, memory, knowledge, relationship, reaching out, self, media and tools, games, stories, roleplay, operations), each with the same headings.
- An API reference generated from the endpoints by a script, so it can’t drift.
The cleanup plan ended with a documentation sweep, but the real rule is that every change updates the docs in the same commit.
What I learned
- Pay down structure before it’s urgent. 51 dependencies on one class is a smell long before it’s a crisis. A phased, behaviour-preserving cleanup cost a fraction of what a rewrite would have.
- Seams are worth more than layers. The four projects keep dependencies pointing the right way. The single seam per system is what makes things actually replaceable.
- Make architecture rules tests. And check that each test can fail.
- Fake the model, not the database. Real PostgreSQL and Qdrant with scripted models gave tests that are fast, deterministic, and still catch real bugs.
- Write docs for the system, not just for people. The changelog and glossary being part of Aria’s knowledge is the best reason I’ve found to keep them current.
One post left: Extras, the two features that didn’t fit anywhere else.