Architecture
The query path
Eight stages#
Resolve the tenant, once
TenantScopeA scope is constructed at the top of the request and threaded through everything below it. Functions take aTenantScope, never atenant_id: str— so there is no signature in the retrieval layer that would even accept an unscoped read.Rewrite a follow-up into a standalone question
skipped on turn 1Skipped entirely when there is no history: a question with no conversation behind it is standalone by definition, and rewriting could only paraphrase it at the cost of a round trip and a chance to mangle an error code. Otherwise the last six turns go in at temperature 0.0, and the output is validated before it is used.Probe the cache — exact first, then semantic
after the rewriteExact is an O(1) lookup that cannot be wrong. Semantic needs an embedding and a scan and can be wrong, so the cheap safe one goes first. Identifier-bearing queries skip the semantic path entirely, on read and on write.Retrieve
~29 ms p50The keyword and vector legs run concurrently against different indexes on separate pooled connections, so hybrid retrieval costs roughly max(legs) rather than their sum. Their ranked lists are merged by reciprocal rank fusion into twenty candidates.Rerank the top eight
1.7–3.3 s on CPUA local cross-encoder scores query–chunk pairs and emits the top five. Only eight candidates reach it: cost is linear in pairs, and narrowing the second stage rather than the first means recall@20 still measures the same twenty chunks.Gate on confidence, before generation
two thresholdsBelow the threshold, the pipeline abstains and opens an escalation without calling the model at all. Which threshold applies is read from the data — whether a reranker score is present — not from configuration, because configuration knows the intent and only the result knows what happened.Generate, closed-book, streamed
temperature 0.1The model sees the retrieved chunks and nothing else, with a few-shot prompt that requires an inline citation per claim and instructs it to prefer the newest source and flag a discrepancy when sources conflict. Not temperature 0.0 — open models repeat themselves there.Validate the citations, then record
~30 msThe answer is split into sentence-level claims and each is checked against the chunk it cited, in one batched embedding call. Then the trace row is written and the answer is cached — unless it was an abstention, which is never cached.
The streaming contract#
Three server-sent event types arrive in a fixed order, and the ordering is a designed feature rather than an implementation detail.
| Event | When | Payload |
|---|---|---|
meta | Before the first token of the answer | Numbered citations, the gate decision, the rewrite result, cache status |
delta | Repeatedly, as the model produces text | One fragment of the answer |
final | Once, last | The complete response: timings, cost, provider, trace id, citation report |
error | Only after the stream has already begun | A terminal message — the HTTP status is long gone by then |
The client parses SSE by hand, and must
The browser has a built-in SSE client, EventSource, and it is GET-only. A chat request carries a tenant, a query and the conversation history — far too much for a query string, and putting a conversation in a URL is a bad idea regardless, because it lands in logs, in history and in referrers.
So the frontend uses fetch with a POST body and reads the response stream. The rule that makes it correct: a network chunk is not an event. One read can deliver half an event, three events, or an event split mid-JSON, so bytes accumulate into a buffer and only complete \n\n-separated blocks are parsed, with the remainder kept for the next read. The decoder is called with { stream: true } for the same reason: a multi-byte character split across two chunks would otherwise decode as a replacement character mid-word.
Every way out#
There are seven, and six of them are not the happy path. Enumerating them is worth more than describing the successful one, because the whole product thesis is about what happens when the system cannot answer.
┌─ cache hit -> _serve_cached() action = "cache_hit"
│
POST /chat -> stream() ┼─ every leg failed -> _abstain(no_results)
│
├─ gate: no results -> _abstain(no_results)
├─ gate: below threshold-> _abstain(low_confidence)
│
├─ generation failed -> _abstain(generation_failed)
│ (+ an "error" event if text had already gone out)
│
├─ the MODEL abstained -> escalation, action = "escalated"
│
└─ success -> validate -> cache -> action = "answered"Three abstention paths, one exit
The gate abstains on weak retrieval. The model abstains when the context was topically right and factually silent. Generation failure abstains when every provider died. All three route through one _abstain() function, all three write an escalation row, and all three set action = "escalated" on the trace.
The reason is metric integrity. Three code paths setting the escalation row, the trace action and the user-facing sentence independently is exactly how an escalation-rate metric ends up lying about the system it measures. And an abstention is streamed as meta then delta like any other answer, not as an error shape — otherwise every consumer would need two branches for one legitimate outcome.
| Abstention path | What it actually means | Where the fix lives |
|---|---|---|
| Gate, below threshold | Retrieval found something, and not strongly enough to answer from | Retrieval, chunking, or the threshold itself |
| Gate, no results | Nothing matched at all — usually genuinely out of scope | The corpus, if the question keeps recurring |
| Model declined | The context was on-topic and did not contain the answer | The corpus. This is the most interesting signal of the three — it is a documentation gap, not a retrieval bug |
| Generation failed | Every provider in the chain errored or was exhausted | Provider configuration or quota. A spike here beside a flat low-confidence count is a much clearer outage signal than a raw 500 count |
What happens when each dependency fails#
| When this fails | Behaviour | Where you see it |
|---|---|---|
| One retrieval leg | The other leg's results are used alone; the request continues | degraded_legs on the answer's timing strip |
| Both retrieval legs | Abstain with no_results, and open an escalation | Escalation banner, and the reason on the dashboard |
| Redis | Every cache call is wrapped, so a failure degrades to a miss | Cache hit rate falling to zero |
| Query rewriting | Fall back to the raw query and record why. A degraded rewrite gives worse retrieval; a raised exception gives no answer at all | skipped_reason on the rewrite result |
| One LLM provider | Fail over to the next in the chain. A Retry-After larger than the maximum wait is read as quota exhaustion and fails over immediately rather than sleeping | Provider usage on the dashboard |
| Every LLM provider | Abstain with generation_failed | A spike in that escalation reason |
| Citation validation | The answer is still returned, with no report rather than no answer | An empty citation report |
| Writing the trace | Swallowed. Observability failure must never become a user-facing one | Server logs only |