Skip to content

Reference

Glossary

Terms this documentation uses in a specific sense. A term earns a place here only if using it loosely would cause a misreading somewhere else on the site.

Retrieval#

Leg
One of the two independent retrieval paths — keyword or vector. A leg supplies a query fragment and never writes a full query; the tenant scope composes the SQL around it.
Candidates vs results
Candidates is the fused list before reranking; results is what the cross-encoder emitted. Both are kept on every retrieval, because scoring the wrong one made a working feature look worthless once already.
Reciprocal rank fusion
Merging two ranked lists by Σ weight / (k + rank) rather than by score. Scale-free, so there is no normalisation step to get wrong.
Cross-encoder
A model that scores a query and a chunk jointly rather than comparing two independently-computed vectors. More accurate, and impossible to precompute — which is why it costs seconds on a CPU.
Margin
(s₁ − s₅) / s₁ over fused scores, used to decide whether reranking is worth running. On this corpus it turned out to measure “was the top five unanimous”, not “how confident is the top result”.
Degraded leg
A retrieval leg that errored. The request continues on the other leg's results, and the answer says so — it was built on less evidence than usual.

Scores and gating#

Score kind
Which of the two scales a confidence figure is on: rerank (a sigmoid in [0, 1]) or fused (a fusion score around 0.016–0.033). They differ by roughly 30x, so a bare number is uninterpretable. Recorded on every trace and shown on every confidence pill.
Confidence gate
The check that runs BEFORE generation. Below threshold, the pipeline abstains without calling the model at all — which is what makes an out-of-scope question cost zero.
Abstention
A deliberate refusal to answer, using a fixed sentence that the prompt, the detector and the eval assertions all agree on. It is the system working, and it is never rendered as an error.
Escalation
The row written whenever the system abstains, carrying the conversation and the top ten retrieved chunks with their scores. Ten rather than five, because 'the right chunk was at rank 7' and 'the right chunk was never retrieved' are different bugs.

Content and versioning#

Chunk
The unit of retrieval. Sized by source type — a docs section, a changelog entry, or one ticket question-and-resolution pair — never by a fixed character window.
Heading path
A docs chunk's position in the document hierarchy, e.g. Billing > Invoices > Proration. Prepended into the chunk's content so it reaches both the keyword index and the embedding, and stored separately for display.
Source locator
How the golden set names a source — a slug and heading, or an entry or ticket id — rather than a chunk UUID. Resolved to chunk ids at the start of every run, which is what makes the chunking experiment possible at all.
Superseded
A document a changelog entry explicitly retired. Its chunks are marked not-current and become unretrievable, but are never deleted.
Contested
A chunk a newer changelog entry contradicts on one fact, while the rest of the page stays correct. Both stay live; the chunk is tagged at ingestion, and the UI shows an amber “superseded in part” badge. This is the realistic case — in production nobody remembers to mark the old doc.
Shadow tenant
A throwaway tenant used to ingest the same corpus a second way, so two chunking strategies can be scored side by side against the same golden set.

Generation and validation#

Closed-book
The model sees the retrieved chunks and nothing else, and is instructed to abstain rather than answer from its own knowledge.
Claim
One sentence of the answer, treated as the unit of validation. A claim citing two sources is supported if EITHER backs it, because the prompt explicitly asks the model to cite both sides of a conflict.
Grounding rate
The fraction of an answer's claims that matched the source they cited. Precomputed onto the trace row so the dashboard and the triage script cannot disagree about it.
Fabricated citation
A [n] marker pointing at a source index the backend never offered. Flagged in rose in the answer text rather than suppressed — the answer may still be correct, and hiding the discrepancy would be its own failure.
Query rewriting
Turning a follow-up into a standalone question using recent history. Skipped entirely on the first turn, and its output is validated — the dominant failure of a rewriting prompt is that the model answers the question instead.

Caching and cost#

Exact hit vs semantic hit
An exact hit is the same question asked again. A semantic hit is a DIFFERENT question that embedded close enough — which the reader deserves to know while judging whether the answer fits.
Identifier guard
The rule that keeps queries containing error codes, versions, status codes, ticket ids or endpoint paths out of the semantic cache — on write as well as read. Two error codes mean nearly the same thing and have opposite answers.
Active invalidation
Deleting exactly the cached answers built on chunks a re-ingest changed, via a chunk-to-keys reverse index — rather than wiping the tenant's cache or waiting for a TTL.
Virtual cost
What the observed token usage would cost at paid-API list prices. Real spend is $0 on free tiers. The label travels with the figure everywhere, because detached from it the number reads as a bill.

Evaluation#

Golden set
65 cases across six types, derived from the corpus specification and hand-edited. Ground truth is only trustworthy because every fact in the corpus was declared in code rather than generated.
Arm
One configuration under comparison — 'vector only', 'BM25 + vector + reranker'. Two arms are only comparable if they were scored over the same cases and the same list.
Hard assertion
A binary correctness check — must-abstain, no cross-tenant leak, no fabricated citations — with zero tolerance, kept separate from quality metrics so a 5% quality budget cannot absorb a security bug.
Macro-averaging
Taking the mean over CASES rather than over retrieved chunks, so one case expecting six chunks cannot dominate fifty expecting one. Micro-averaging would let the corpus shape drive the headline number.
Baseline
A committed scorecard the CI gate compares against. Committed rather than taken from the previous run, because auto-updating lets quality erode one tolerated 4% drop at a time.