Skip to content

Architecture

Ingestion and chunking

Three source types, three chunking strategies, and a second pass that runs only after every document has landed. Chunking is the highest-leverage decision in the whole system: it beat every retrieval-side change that was measured against it.

The pipeline#

  1. Load

    3 loaders
    Product docs are Markdown with YAML frontmatter; the changelog and the ticket archive are JSONL, one record per line. Each loader emits the same ParsedDocument, so nothing downstream branches on where a document came from except the chunker selection.
  2. Deduplicate by content hash

    idempotent
    A document whose content_hash already exists is skipped entirely. This is what makes re-running ingestion nearly free, and it is why the quickstart can tell you to just run it again if you are not sure whether it finished.
  3. Chunk, by source type

    outside the transaction
    Three strategies, described below. Chunking is CPU work and embedding is model work, so both happen before the transaction opens — holding a database transaction across two minutes of embedding would be the easiest way to turn a slow ingest into a lock-contention incident.
  4. Embed, through a cache

    sha256(model + text)
    Only cache misses reach the model, in batches of 32. The cache is keyed on the model name as well as the text, because embeddings from different models live in different spaces and are never comparable — a key without the model name would silently mix two vector spaces in one index.
  5. Insert transactionally

    document + chunks
    The document row and all of its chunks land in one transaction, and any previous version of that document is archived — marked is_current = false, never deleted. Chunk ids for the outgoing version are collected before the delete, because afterwards there is no way to know which cached answers went stale.
  6. Second pass, after every document

    supersede + tag
    Changelog entries can retire a document or contest one fact on it. Both are applied only once the whole corpus is in.
  7. Invalidate the cache, after the commit

    reverse index
    Every cached answer records which chunks it was built from, so re-ingesting a document deletes precisely the answers that depended on it. Running this after the commit is deliberate: inside the transaction a rollback would clear the cache for content that still exists, which costs a few extra LLM calls. The reverse — new content committed while stale answers survive — is much worse.

Three chunking strategies#

Each source type has a natural unit, and a fixed-size window destroys all three of them. A docs page has sections; a changelog has entries; a ticket has a question and its resolution, which are worthless apart.

SourceUnitTarget sizeSpecial handling
Product docsOne heading section300–500 tokHeading path prepended into the content; 15% overlap; tables and code fences never split; stub sections merged with siblings
ChangelogOne entry50–200 tokVersion and date appear both in the text — for the keyword leg and the conflict rule — and in metadata, for recency
Support ticketsOne question–resolution pair≤ 500 tokError code repeated in the header line; a long ticket splits at the question/answer seam and each half is re-labelled

The heading path goes into the content, not just the metadata

A docs chunk's stored content begins with its position in the document hierarchy:

text
content = "Billing > Invoices > Proration\n\n<section body>"

This is decided at ingestion time and is expensive to change later, because both the keyword index and the embedding are computed from content — reversing it means re-chunking and re-embedding everything.

The reason is that a section's body frequently never repeats its own topic words. “Proration is calculated daily…” sitting under the heading Webhooks > Retry Logic would be unfindable for the query “webhook retry”. Prepending puts those words into both the text-search vector and the embedding at once, and gives the generator the chunk's position in the document without extra prompt assembly. The clean path is also stored in its own column, for display and metadata filtering.

Token budgets are measured with the model's own tokenizer

A chunk that is “450 tokens” by word count can be 700 real tokens. It gets stored whole, embedded truncated at the model's 512 limit, and retrieval quietly degrades with nothing in any log. Silent truncation is the worst class of ingestion bug because everything appears to work.

So budgets are measured with the embedding model's HuggingFace tokenizer. And because per-unit counts are not additive — joining adds separators, and subword merging behaves differently across a boundary — the docs chunker measures the actual joined candidate string and runs a final enforcement pass over finished chunks. Summing the parts underestimates, which is precisely how a chunk sneaks past the cap.

Versioning and planted conflicts#

Stale data is the failure mode this system exists to handle, so the corpus contains two kinds of it. A corpus with only one kind can only demonstrate one defence.

KindMechanismIngestion behaviourWhich defence it tests
Supersededsupersedes: <slug>The document and its chunks become is_current = falseIngestion-time metadata discipline
Unmarked conflictconflicts_with: <slug>Both stay live; the doc's chunks are tagged conflicts_with_entryThe generation-time conflict rule, and the amber badge in the UI

Why the contested document is not archived too

Because the changelog contradicts one fact on that page — “the retry limit is now 5” — not the page. The rest of “Webhooks Overview” is still correct, and archiving it would destroy good information to fix one stale sentence.

So the conflict is deliberately pushed to generation time, where the model must prefer the newest source and flag the discrepancy, and where the interface can show the contested source with an amber badge. This is also the realistic case: in production nobody remembers to mark the old document, and a system that only handles declared supersession handles the easy half of the problem.

How the corpus itself is built#

The evaluation harness needs to map queries to known-correct sources, which is only possible if the corpus contents are known with certainty. So the corpus is generated hybrid: every document, version, date, error code and planted conflict is declared in Python, and only the body prose is written by a model, keyed by prompt hash into a committed disk cache with a deterministic template fallback.

ApproachWhy it was rejected
Fully templatedPerfectly reproducible, and the prose is formulaic in a way that makes retrieval unrealistically easy — every chunk reads the same, so lexical overlap with queries is artificial
Fully LLM-generatedRealistic, and the facts drift. A model asked for sixty doc pages invents its own error codes and contradicts itself, so ground truth becomes unverifiable and regeneration silently changes it
Hybrid — declared facts, generated proseFacts the evaluation depends on are never generated. Prose, which only needs to be plausible, is. The cache makes reruns byte-identical, and the fallback means a fresh clone works with no API key at all
A guardrail warns — rather than fails — when a required literal from the brief does not survive into the generated text. A hundred-document run should not abort, but you must know before building a golden set on it.