Aria's knowledge, Part 1: Reading the world safely
The second RAG pipeline: Aria reads feeds, sitemaps, pages and JSON APIs in the background, fetches them without ever touching your private network, splits and indexes articles, and cites its sources in replies, with citations that come from the database, never from the model.
The introduction described Aria’s three kinds of memory: what it knows about you, what it knows about the world, and what it learns you care about. Memory covered the first. The reaching-out posts mentioned “news on topics you follow” in passing. This post is the second kind: world knowledge, and how Aria gets it.
It’s the same RAG pattern as memory (Part 2 of the memory posts), with one big difference: the input is the open web. That makes it a security problem first and a retrieval problem second.
Kept apart from memory
World knowledge is completely separate from personal memory: its own tables in PostgreSQL (sources, documents, document versions, chunks, usage events), and its own Qdrant collection. A news article can never be mistaken for something you said, and deleting all your memories never touches the library.
As with memory, PostgreSQL is the source of truth and the vectors can be rebuilt from it.
Sources
You add sources on the Knowledge page. Four kinds:
- RSS, Atom or RDF feeds.
- Sitemaps (including sitemap indexes).
- Web pages, read as a single document and re-read on every poll.
- JSON APIs, with a small field mapping: where the items are, which field is the URL, and optionally title, summary, date and author, each as a dot path.
Single URLs can also be added directly (“read this”). Each source has a schedule, a priority, include and exclude patterns, a limit per poll, and a reading mode: read everything, or read only what’s interesting (Part 2).
The pipeline
Everything runs as background jobs on the same PostgreSQL queue as the rest of Aria:
- Discovery. A scheduler checks every minute which sources are due. Discovery fetches the feed or sitemap (with conditional requests, so an unchanged feed costs almost nothing), normalises the URLs (tracking parameters, fragments and default ports removed), applies the patterns and limits, and stores new URLs as discovered. A source that keeps failing backs off, up to a day.
- Fetch and extract. Each article is fetched and run through a Readability-style extractor into structured text: headings, paragraphs and list items. A new version is stored only when the normalised content changed. The raw HTML is never stored. A page with the same canonical URL or identical content as another document is rejected as a duplicate.
- Analyse and index. The analysis model reads the article and returns topics, entities, a summary, a category, keywords and a content type (news, opinion, reference…). Then the text is split into chunks of about 350 tokens, structure-aware, with the headings repeated and a short overlap, embedded with the title, and stored in Qdrant. When a new version arrives, the old version’s chunks are deactivated and their vectors removed.
Rate limits apply on top (KNOWLEDGE_MAX_ARTICLES_PER_HOUR and PER_DAY), so a busy feed can’t flood the model server.
Failures are sorted by kind. A 404, a robots.txt refusal, an oversized or unsupported file is rejected and not retried. A network error or a 5xx is retried with backoff. A page with no extractable text, usually a JavaScript app, is marked as needing a browser. With the browser container switched on, Aria renders the page in a real browser and extracts it again.
Fetching safely
A component that fetches arbitrary URLs from inside your home network is a textbook SSRF risk: a malicious feed could point an “article” at your router’s admin page, a NAS, or a cloud metadata endpoint. Aria’s fetcher was built with that in mind from the start:
- http and https only, no credentials in URLs, no proxy.
- Every connection is checked at connect time, against the actual IP address being connected to: loopback, private ranges, link-local (which includes cloud metadata addresses), carrier-grade NAT, multicast, documentation and reserved ranges, and IPv6 forms that embed any of those. Checking at connect time, not when resolving the name, also defeats DNS rebinding.
- Redirects are followed by hand, at most five, and each one is checked again.
- robots.txt is respected (RFC 9309, cached for an hour, crawl-delay honoured), with a delay per host and a global limit on concurrent fetches.
- Timeouts, size limits and a content-type allow-list.
- XML parsing ignores DTDs and never resolves external entities, which closes the XXE door in feeds and sitemaps.
- JSON is parsed with a depth limit, and only the mapped fields are read.
The same address policy is used by Aria’s browser tool, which the tools posts cover: it can only ever reach public addresses.
Retrieval
When you write, Aria searches the library with your message and the previous one as the query. Candidates must pass a minimum similarity first, because relevance to the conversation is what matters most. Then they’re ranked:
score = similarity × 0.60 + topic alignment × 0.15 + freshness × 0.10
+ source weight × 0.10 + your interest × 0.05
Two details make it work better than plain similarity:
- Freshness depends on content type. News halves in value every 3 days, time-sensitive content every 30, opinion every 90. Reference and evergreen content doesn’t decay. Age uses the publication date, not when Aria fetched it. When you ask about recent developments, freshness counts double.
- Recently used documents get a small penalty, so Aria doesn’t cite the same article in every reply.
At most two excerpts per document are used. Knowledge is also the first thing dropped when the prompt is too big (Conversations, Part 3): nice to have, never essential.
Citing sources
Excerpts go into the prompt as reference data, each with a label ([K1], [K2]), its title, source, publication date and URL, with instructions to attribute claims and keep disagreeing sources apart. When excerpts from different sources cover the same story, they carry a story label so Aria can compare them instead of blending them.
The important rule: citation metadata always comes from the database, never from the model. The chat shows the sources under a reply from the records of what was actually retrieved and used, not from whatever the model wrote. A model can’t invent a citation, misattribute a quote to the wrong site, or get a URL slightly wrong, because it never produces the citation in the first place.
Every retrieved excerpt is also recorded as used in a prompt or dropped for space, which is what drives the “recently used” penalty.
Its own documentation
One source is special. Aria’s own development history, the document I update with every change, is bundled with the app and loaded into the library at startup as “Aria’s own documentation”, along with the glossary of its systems. It’s re-indexed only when the text changes. It’s retrieved like any other article, so when you ask Aria what it can do or what changed recently, it answers from its own docs and says so. Knowing itself builds on that.
What I learned
- Fetching the web from home is a security feature first. Checking the address at connect time, re-checking every redirect and ignoring DTDs are the things that make a background reader safe to run on a home network.
- Same pattern, separate store. Reusing the memory pipeline’s design (database first, rebuildable index, ranked retrieval, data labels) made knowledge quick to build. Keeping it in its own collection kept it from contaminating memory.
- Freshness is per content type. A single decay curve either buries reference material or keeps stale news.
- Never let the model write citations. Sources from the retrieval records are always right. Sources from the model are right most of the time.
In Part 2: how Aria decides what to read. Interests learned from conversation, an adaptive reader that balances them with exploration, and stories, entities and trends across sources.