Skip to content

Architecture

The query path

One readable sequence in app/generation/pipeline.py. Each stage may decline to hand work to the next, and every way of declining routes through a single exit — because an escalation row, a trace's action and the sentence the user reads must never be able to disagree.

Eight stages#

  1. Resolve the tenant, once

    TenantScope
    A scope is constructed at the top of the request and threaded through everything below it. Functions take a TenantScope, never a tenant_id: str — so there is no signature in the retrieval layer that would even accept an unscoped read.
  2. Rewrite a follow-up into a standalone question

    skipped on turn 1
    Skipped entirely when there is no history: a question with no conversation behind it is standalone by definition, and rewriting could only paraphrase it at the cost of a round trip and a chance to mangle an error code. Otherwise the last six turns go in at temperature 0.0, and the output is validated before it is used.
  3. Probe the cache — exact first, then semantic

    after the rewrite
    Exact is an O(1) lookup that cannot be wrong. Semantic needs an embedding and a scan and can be wrong, so the cheap safe one goes first. Identifier-bearing queries skip the semantic path entirely, on read and on write.
  4. Retrieve

    ~29 ms p50
    The keyword and vector legs run concurrently against different indexes on separate pooled connections, so hybrid retrieval costs roughly max(legs) rather than their sum. Their ranked lists are merged by reciprocal rank fusion into twenty candidates.
  5. Rerank the top eight

    1.7–3.3 s on CPU
    A local cross-encoder scores query–chunk pairs and emits the top five. Only eight candidates reach it: cost is linear in pairs, and narrowing the second stage rather than the first means recall@20 still measures the same twenty chunks.
  6. Gate on confidence, before generation

    two thresholds
    Below the threshold, the pipeline abstains and opens an escalation without calling the model at all. Which threshold applies is read from the data — whether a reranker score is present — not from configuration, because configuration knows the intent and only the result knows what happened.
  7. Generate, closed-book, streamed

    temperature 0.1
    The model sees the retrieved chunks and nothing else, with a few-shot prompt that requires an inline citation per claim and instructs it to prefer the newest source and flag a discrepancy when sources conflict. Not temperature 0.0 — open models repeat themselves there.
  8. Validate the citations, then record

    ~30 ms
    The answer is split into sentence-level claims and each is checked against the chunk it cited, in one batched embedding call. Then the trace row is written and the answer is cached — unless it was an abstention, which is never cached.

The streaming contract#

Three server-sent event types arrive in a fixed order, and the ordering is a designed feature rather than an implementation detail.

EventWhenPayload
metaBefore the first token of the answerNumbered citations, the gate decision, the rewrite result, cache status
deltaRepeatedly, as the model produces textOne fragment of the answer
finalOnce, lastThe complete response: timings, cost, provider, trace id, citation report
errorOnly after the stream has already begunA terminal message — the HTTP status is long gone by then

The client parses SSE by hand, and must

The browser has a built-in SSE client, EventSource, and it is GET-only. A chat request carries a tenant, a query and the conversation history — far too much for a query string, and putting a conversation in a URL is a bad idea regardless, because it lands in logs, in history and in referrers.

So the frontend uses fetch with a POST body and reads the response stream. The rule that makes it correct: a network chunk is not an event. One read can deliver half an event, three events, or an event split mid-JSON, so bytes accumulate into a buffer and only complete \n\n-separated blocks are parsed, with the remainder kept for the next read. The decoder is called with { stream: true } for the same reason: a multi-byte character split across two chunks would otherwise decode as a replacement character mid-word.

Every way out#

There are seven, and six of them are not the happy path. Enumerating them is worth more than describing the successful one, because the whole product thesis is about what happens when the system cannot answer.

app/generation/pipeline.py — exit paths
                       ┌─ cache hit           -> _serve_cached()  action = "cache_hit"
                       │
POST /chat -> stream() ┼─ every leg failed     -> _abstain(no_results)
                       │
                       ├─ gate: no results     -> _abstain(no_results)
                       ├─ gate: below threshold-> _abstain(low_confidence)
                       │
                       ├─ generation failed    -> _abstain(generation_failed)
                       │                          (+ an "error" event if text had already gone out)
                       │
                       ├─ the MODEL abstained  -> escalation, action = "escalated"
                       │
                       └─ success              -> validate -> cache -> action = "answered"

Three abstention paths, one exit

The gate abstains on weak retrieval. The model abstains when the context was topically right and factually silent. Generation failure abstains when every provider died. All three route through one _abstain() function, all three write an escalation row, and all three set action = "escalated" on the trace.

The reason is metric integrity. Three code paths setting the escalation row, the trace action and the user-facing sentence independently is exactly how an escalation-rate metric ends up lying about the system it measures. And an abstention is streamed as meta then delta like any other answer, not as an error shape — otherwise every consumer would need two branches for one legitimate outcome.

Abstention pathWhat it actually meansWhere the fix lives
Gate, below thresholdRetrieval found something, and not strongly enough to answer fromRetrieval, chunking, or the threshold itself
Gate, no resultsNothing matched at all — usually genuinely out of scopeThe corpus, if the question keeps recurring
Model declinedThe context was on-topic and did not contain the answerThe corpus. This is the most interesting signal of the three — it is a documentation gap, not a retrieval bug
Generation failedEvery provider in the chain errored or was exhaustedProvider configuration or quota. A spike here beside a flat low-confidence count is a much clearer outage signal than a raw 500 count

What happens when each dependency fails#

When this failsBehaviourWhere you see it
One retrieval legThe other leg's results are used alone; the request continuesdegraded_legs on the answer's timing strip
Both retrieval legsAbstain with no_results, and open an escalationEscalation banner, and the reason on the dashboard
RedisEvery cache call is wrapped, so a failure degrades to a missCache hit rate falling to zero
Query rewritingFall back to the raw query and record why. A degraded rewrite gives worse retrieval; a raised exception gives no answer at allskipped_reason on the rewrite result
One LLM providerFail over to the next in the chain. A Retry-After larger than the maximum wait is read as quota exhaustion and fails over immediately rather than sleepingProvider usage on the dashboard
Every LLM providerAbstain with generation_failedA spike in that escalation reason
Citation validationThe answer is still returned, with no report rather than no answerAn empty citation report
Writing the traceSwallowed. Observability failure must never become a user-facing oneServer logs only

Retrieval and ranking, in detail →