Architecture
Generation and citations
The confidence gate#
The gate sits before generation, which is what makes an out-of-scope question cost zero LLM calls rather than one wasted one. It reads the top retrieved result's score and compares it against a threshold.
The complication is that there are two thresholds, because there are two score scales roughly 30x apart: the reranker's sigmoid in [0, 1] and the fusion score around 0.016–0.033.
| Threshold | Value | Applies when |
|---|---|---|
confidence_threshold_rerank | 0.45 | A reranker score is present on the top result |
confidence_threshold_fused | 0.015 | Reranking did not run — the score is a fusion score |
A single threshold would be calibrated for at most one scale and silently wrong for the other. Set for the reranker it would abstain on literally everything unreranked; set for fusion it would pass everything reranked. Both failures present as “the gate is not working” and neither points at the cause.
Query rewriting#
A follow-up like “what about the backoff?” is not a searchable query. Rewriting turns it into a standalone question using the last six turns, at temperature 0.0.
Four rules that make it safe
| Rule | Why |
|---|---|
| Skipped entirely on the first turn | A question with no conversation behind it is standalone by definition. Rewriting could only paraphrase, at the cost of a round trip and a chance to mangle an error code — and most support sessions are a single turn, so this is not a micro-optimisation |
| The output is validated before use | The dominant failure of a rewriting prompt is that the model ANSWERS the question instead of rewriting it. An answer silently substituted for a query means searching the corpus for the text of a hallucinated response — a spectacular retrieval bug that raises nothing |
| History is rendered as data, not replayed as turns | We want the model analysing the conversation, not participating in it. Replayed as real turns it reliably answers instead of rewriting |
| The prompt protects identifiers explicitly | A rewriter that helpfully expands ERR_TIMEOUT_502 into 'timeout error' destroys the exact-match signal the keyword leg exists for. Pinned by a test |
The validator rejects empty output, output over 400 characters, output that grew more than 6x and is long in absolute terms, multi-paragraph output, and text opening with a refusal or “Sure, here is”. On any rejection the raw query is used and the reason recorded — a degraded rewrite gives worse retrieval, whereas a raised exception gives no answer at all.
The generation prompt#
Closed-book: the model sees the retrieved chunks and nothing else, with a numbered context block, a few-shot example, and a small set of rules.
1. Answer ONLY from the numbered sources below. If they do not contain the
answer, say the fixed abstention sentence and nothing else.
2. Every factual claim carries an inline [n] marker naming the source it came from.
3. Do not merge two sources into one claim without citing both.
4. If sources disagree, prefer the one with the newer version or effective date,
AND say that they disagree.
5. Never invent an error code, a version number, or a limit.| Setting | Value | Why |
|---|---|---|
llm_temperature | 0.1 | Factual answers want 0.0–0.2. Not 0.0, because open models repeat themselves there |
query_rewrite_history_turns | 6 | Three exchanges. More drags stale entities from an earlier, unrelated problem into the query — the subtler of the two failures |
abstention_message | fixed string | The prompt, the detector that notices the model abstained, and the eval assertions must all agree on the exact sentence |
Citation validation#
After the answer exists, it is split into sentence-level claims. Each claim is embedded with the same local encoder retrieval uses and compared against the chunks it cited. Threshold 0.50. All claims and all cited chunks go into one batched embedding call, so validation latency does not scale with answer length — otherwise a more helpful answer would be a slower one, which is a perverse incentive.
| Option | Verdict |
|---|---|
| Embedding similarity | Free, local, deterministic, ~30 ms. Runs on EVERY answer, which is the property that matters — a check that runs always beats a better check that gets disabled for being slow |
| LLM-as-judge | Much stronger — catches contradiction, not just topical drift. Costs a call per answer, adds seconds, burns the quota the evaluation needs, and introduces a second model whose failures correlate with the generator's |
| An entailment model | The technically right answer, and another ~1.4 GB model on a CPU already spending seconds on reranking. The latency would land on every answer |
A claim citing two sources needs only one to support it
Because rule 4 of the prompt explicitly asks the model to cite both sides of a conflict. Requiring every cited source to support the claim would punish the exact behaviour that was requested.
What the interface does with the result#
| Signal | How it appears |
|---|---|
| Claim supported | Emerald verdict inside the expanded source, with the similarity |
| Claim unsupported | Rose verdict, and a count in the panel header |
| Marker pointing at a source that was never offered | Rendered in rose in the answer text and listed as fabrication — flagged, not suppressed. The answer may still be correct, and hiding the discrepancy would be its own failure |
| Source retrieved but never cited | Dimmed rather than hidden — in aggregate it is a retrieval-precision signal |
| Chunk contested by a newer entry | Amber 'superseded in part' badge, computed at ingestion time rather than guessed in the browser |
What happens to the finished answer — and the one kind that is never cached →