Architecture
Ingestion and chunking
The pipeline#
Load
3 loadersProduct docs are Markdown with YAML frontmatter; the changelog and the ticket archive are JSONL, one record per line. Each loader emits the sameParsedDocument, so nothing downstream branches on where a document came from except the chunker selection.Deduplicate by content hash
idempotentA document whosecontent_hashalready exists is skipped entirely. This is what makes re-running ingestion nearly free, and it is why the quickstart can tell you to just run it again if you are not sure whether it finished.Chunk, by source type
outside the transactionThree strategies, described below. Chunking is CPU work and embedding is model work, so both happen before the transaction opens — holding a database transaction across two minutes of embedding would be the easiest way to turn a slow ingest into a lock-contention incident.Embed, through a cache
sha256(model + text)Only cache misses reach the model, in batches of 32. The cache is keyed on the model name as well as the text, because embeddings from different models live in different spaces and are never comparable — a key without the model name would silently mix two vector spaces in one index.Insert transactionally
document + chunksThe document row and all of its chunks land in one transaction, and any previous version of that document is archived — markedis_current = false, never deleted. Chunk ids for the outgoing version are collected before the delete, because afterwards there is no way to know which cached answers went stale.Second pass, after every document
supersede + tagChangelog entries can retire a document or contest one fact on it. Both are applied only once the whole corpus is in.Invalidate the cache, after the commit
reverse indexEvery cached answer records which chunks it was built from, so re-ingesting a document deletes precisely the answers that depended on it. Running this after the commit is deliberate: inside the transaction a rollback would clear the cache for content that still exists, which costs a few extra LLM calls. The reverse — new content committed while stale answers survive — is much worse.
Three chunking strategies#
Each source type has a natural unit, and a fixed-size window destroys all three of them. A docs page has sections; a changelog has entries; a ticket has a question and its resolution, which are worthless apart.
| Source | Unit | Target size | Special handling |
|---|---|---|---|
| Product docs | One heading section | 300–500 tok | Heading path prepended into the content; 15% overlap; tables and code fences never split; stub sections merged with siblings |
| Changelog | One entry | 50–200 tok | Version and date appear both in the text — for the keyword leg and the conflict rule — and in metadata, for recency |
| Support tickets | One question–resolution pair | ≤ 500 tok | Error code repeated in the header line; a long ticket splits at the question/answer seam and each half is re-labelled |
The heading path goes into the content, not just the metadata
A docs chunk's stored content begins with its position in the document hierarchy:
content = "Billing > Invoices > Proration\n\n<section body>"This is decided at ingestion time and is expensive to change later, because both the keyword index and the embedding are computed from content — reversing it means re-chunking and re-embedding everything.
The reason is that a section's body frequently never repeats its own topic words. “Proration is calculated daily…” sitting under the heading Webhooks > Retry Logic would be unfindable for the query “webhook retry”. Prepending puts those words into both the text-search vector and the embedding at once, and gives the generator the chunk's position in the document without extra prompt assembly. The clean path is also stored in its own column, for display and metadata filtering.
Token budgets are measured with the model's own tokenizer
A chunk that is “450 tokens” by word count can be 700 real tokens. It gets stored whole, embedded truncated at the model's 512 limit, and retrieval quietly degrades with nothing in any log. Silent truncation is the worst class of ingestion bug because everything appears to work.
So budgets are measured with the embedding model's HuggingFace tokenizer. And because per-unit counts are not additive — joining adds separators, and subword merging behaves differently across a boundary — the docs chunker measures the actual joined candidate string and runs a final enforcement pass over finished chunks. Summing the parts underestimates, which is precisely how a chunk sneaks past the cap.
Versioning and planted conflicts#
Stale data is the failure mode this system exists to handle, so the corpus contains two kinds of it. A corpus with only one kind can only demonstrate one defence.
| Kind | Mechanism | Ingestion behaviour | Which defence it tests |
|---|---|---|---|
| Superseded | supersedes: <slug> | The document and its chunks become is_current = false | Ingestion-time metadata discipline |
| Unmarked conflict | conflicts_with: <slug> | Both stay live; the doc's chunks are tagged conflicts_with_entry | The generation-time conflict rule, and the amber badge in the UI |
Why the contested document is not archived too
Because the changelog contradicts one fact on that page — “the retry limit is now 5” — not the page. The rest of “Webhooks Overview” is still correct, and archiving it would destroy good information to fix one stale sentence.
So the conflict is deliberately pushed to generation time, where the model must prefer the newest source and flag the discrepancy, and where the interface can show the contested source with an amber badge. This is also the realistic case: in production nobody remembers to mark the old document, and a system that only handles declared supersession handles the easy half of the problem.
How the corpus itself is built#
The evaluation harness needs to map queries to known-correct sources, which is only possible if the corpus contents are known with certainty. So the corpus is generated hybrid: every document, version, date, error code and planted conflict is declared in Python, and only the body prose is written by a model, keyed by prompt hash into a committed disk cache with a deterministic template fallback.
| Approach | Why it was rejected |
|---|---|
| Fully templated | Perfectly reproducible, and the prose is formulaic in a way that makes retrieval unrealistically easy — every chunk reads the same, so lexical overlap with queries is artificial |
| Fully LLM-generated | Realistic, and the facts drift. A model asked for sixty doc pages invents its own error codes and contradicts itself, so ground truth becomes unverifiable and regeneration silently changes it |
| Hybrid — declared facts, generated prose | Facts the evaluation depends on are never generated. Prose, which only needs to be plausible, is. The cache makes reruns byte-identical, and the fallback means a fresh clone works with no API key at all |