Skip to content

Architecture

Generation and citations

A gate that decides whether to call the model at all, a closed-book prompt that must cite, and a post-hoc validator that checks every claim against the source it named — with the hole in that check stated at every call site rather than buried in a docstring.

The confidence gate#

The gate sits before generation, which is what makes an out-of-scope question cost zero LLM calls rather than one wasted one. It reads the top retrieved result's score and compares it against a threshold.

The complication is that there are two thresholds, because there are two score scales roughly 30x apart: the reranker's sigmoid in [0, 1] and the fusion score around 0.016–0.033.

ThresholdValueApplies when
confidence_threshold_rerank0.45A reranker score is present on the top result
confidence_threshold_fused0.015Reranking did not run — the score is a fusion score
Both are provisional and labelled so in config. `make tune` sweeps them against the golden set.

A single threshold would be calibrated for at most one scale and silently wrong for the other. Set for the reranker it would abstain on literally everything unreranked; set for fusion it would pass everything reranked. Both failures present as “the gate is not working” and neither points at the cause.

Query rewriting#

A follow-up like “what about the backoff?” is not a searchable query. Rewriting turns it into a standalone question using the last six turns, at temperature 0.0.

Four rules that make it safe

RuleWhy
Skipped entirely on the first turnA question with no conversation behind it is standalone by definition. Rewriting could only paraphrase, at the cost of a round trip and a chance to mangle an error code — and most support sessions are a single turn, so this is not a micro-optimisation
The output is validated before useThe dominant failure of a rewriting prompt is that the model ANSWERS the question instead of rewriting it. An answer silently substituted for a query means searching the corpus for the text of a hallucinated response — a spectacular retrieval bug that raises nothing
History is rendered as data, not replayed as turnsWe want the model analysing the conversation, not participating in it. Replayed as real turns it reliably answers instead of rewriting
The prompt protects identifiers explicitlyA rewriter that helpfully expands ERR_TIMEOUT_502 into 'timeout error' destroys the exact-match signal the keyword leg exists for. Pinned by a test

The validator rejects empty output, output over 400 characters, output that grew more than 6x and is long in absolute terms, multi-paragraph output, and text opening with a refusal or “Sure, here is”. On any rejection the raw query is used and the reason recorded — a degraded rewrite gives worse retrieval, whereas a raised exception gives no answer at all.

The generation prompt#

Closed-book: the model sees the retrieved chunks and nothing else, with a numbered context block, a few-shot example, and a small set of rules.

app/generation/prompts.py (rules, paraphrased)
1. Answer ONLY from the numbered sources below. If they do not contain the
   answer, say the fixed abstention sentence and nothing else.
2. Every factual claim carries an inline [n] marker naming the source it came from.
3. Do not merge two sources into one claim without citing both.
4. If sources disagree, prefer the one with the newer version or effective date,
   AND say that they disagree.
5. Never invent an error code, a version number, or a limit.
SettingValueWhy
llm_temperature0.1Factual answers want 0.0–0.2. Not 0.0, because open models repeat themselves there
query_rewrite_history_turns6Three exchanges. More drags stale entities from an earlier, unrelated problem into the query — the subtler of the two failures
abstention_messagefixed stringThe prompt, the detector that notices the model abstained, and the eval assertions must all agree on the exact sentence

Citation validation#

After the answer exists, it is split into sentence-level claims. Each claim is embedded with the same local encoder retrieval uses and compared against the chunks it cited. Threshold 0.50. All claims and all cited chunks go into one batched embedding call, so validation latency does not scale with answer length — otherwise a more helpful answer would be a slower one, which is a perverse incentive.

OptionVerdict
Embedding similarityFree, local, deterministic, ~30 ms. Runs on EVERY answer, which is the property that matters — a check that runs always beats a better check that gets disabled for being slow
LLM-as-judgeMuch stronger — catches contradiction, not just topical drift. Costs a call per answer, adds seconds, burns the quota the evaluation needs, and introduces a second model whose failures correlate with the generator's
An entailment modelThe technically right answer, and another ~1.4 GB model on a CPU already spending seconds on reranking. The latency would land on every answer

A claim citing two sources needs only one to support it

Because rule 4 of the prompt explicitly asks the model to cite both sides of a conflict. Requiring every cited source to support the claim would punish the exact behaviour that was requested.

What the interface does with the result#

SignalHow it appears
Claim supportedEmerald verdict inside the expanded source, with the similarity
Claim unsupportedRose verdict, and a count in the panel header
Marker pointing at a source that was never offeredRendered in rose in the answer text and listed as fabrication — flagged, not suppressed. The answer may still be correct, and hiding the discrepancy would be its own failure
Source retrieved but never citedDimmed rather than hidden — in aggregate it is a retrieval-precision signal
Chunk contested by a newer entryAmber 'superseded in part' badge, computed at ingestion time rather than guessed in the browser

What happens to the finished answer — and the one kind that is never cached →