Skip to content

Architecture

Retrieval and ranking

Two legs on incomparable score scales, merged by rank rather than by score, then optionally reranked by a cross-encoder that costs more than the entire latency budget. Every part of this was measured, and the measurement disagreed with the design.
The retrieval pipelineA query and a tenant scope enter. The query is embedded, then two legs run concurrently: a BM25 keyword leg using a generated tsvector column with a GIN index and cover-density ranking, and a vector leg using an HNSW index over 384-dimension embeddings. Both legs pass through TenantScope, which welds the tenant predicate onto every read and re-checks every returned row. Their ranked lists are merged by reciprocal rank fusion with k equal to 60 into twenty candidates. The top eight reach a cross-encoder reranker, which emits the top five. The result carries both the reranked results and the pre-rerank candidates, per-leg timings, and any degraded legs.query + scopeembed once, L2-normalisedKeyword legwebsearch_to_tsquery, AND→ORts_rank_cd over a GIN indexVector legSET LOCAL hnsw.ef_search = 100embedding <=> query, HNSWrun concurrently —cost is max(legs)TenantScopeowns FROM chunksWHERE tenant_id = $1AND is_current+ foreign-row tripwireRRF fusionΣ w / (60 + rank) → top 20Cross-encodertop 8 in → sigmoid → top 5conditional rerank: margin ≥ 0.10 skips thisimplemented — and it never fires on this corpusRetrievalResultresults · candidates · per-leg timings · degraded_legsboth lists are kept — see field note 2
The keyword and vector legs run concurrently and neither writes a full query — both hand a fragment to the scope that owns the FROM clause. Fusion needs no hydration step because both legs return full rows.

The keyword leg#

BM25-style keyword matching is provided by Postgres full-text search: a generated tsvector column on chunks with a GIN index, ranked with ts_rank_cd.

The alternative was a real BM25 engine. A Python BM25 library keeps its index in process memory, so it must be rebuilt on every restart and for every worker, and tenant filtering happens in Python after scoring — which is precisely the application-level isolation this system is built to avoid. Elasticsearch is real BM25 and a second datastore to keep consistent with Postgres, which is a large operational bill at this scale.

One datastore for both legs means one tenant predicate enforced in SQL for both, one backup story, and index consistency guaranteed by the generated column.

The query must OR its terms, and every convenient helper ANDs them

This shipped broken and is worth recording honestly, because the failure was silent in exactly the way this project is meant to guard against.

sql
websearch_to_tsquery('english', 'webhook retry limit')
  ->  'webhook' & 'retri' & 'limit'

That is boolean retrieval: a document must contain every term or it does not match at all. BM25 does not work that way — it sums a per-term contribution over the terms a document does contain, so a document matching four of six query terms still scores, just lower. Ranked retrieval is inherently OR-ed; the ranking function decides what wins, not the matcher.

The consequence was severe and quiet. The realistic query "webhook retry limit ERR_TIMEOUT_502" returned zero rows, because the docs page explains retry behaviour and names the error code while the changelog entry is the one that says “limit”. No single chunk held all six lexemes. The keyword leg contributed nothing, fusion degenerated to vector-only, and “hybrid retrieval” became a claim in a README rather than something that happened.

Two operators, two meanings

The naive fix — OR everything — broke exact identifiers instead. Postgres' default parser treats _ as a separator, so ERR_TIMEOUT_502 becomes three independent lexemes and a flat OR matches any chunk containing err. On the real corpus, a query for one error code pulled a completely different error code's ticket into the top five.

The observation that resolves it is that websearch_to_tsquery already distinguishes the two cases:

sql
websearch_to_tsquery('english', 'ERR_TIMEOUT_502')
  ->  'err' <-> 'timeout' <-> '502'        -- followed-by: adjacency IS the point
OperatorWhat it separatesCorrect handling
&Concepts the user typed as separate ideasRewrite to | — this is ranked retrieval
<->One identifier, or a quoted phrase, that the tokenizer happened to splitLeave alone — adjacency is the signal
app/retrieval/bm25.py
SELECT NULLIF(replace(websearch_to_tsquery('english', $2)::text, ' & ', ' | '), '')::tsquery

-- 'webhook' & 'retri' & 'limit' & 'err' <-> 'timeout' <-> '502'
--     becomes
-- 'webhook' | 'retri' | 'limit' | ('err' <-> 'timeout' <-> '502')

Operator precedence does the rest: <-> binds tighter than &, which binds tighter than |, so the phrase group survives intact. A chunk mentioning only err no longer matches the identifier clause; a chunk about webhooks still matches on one concept. NULLIF handles the all-stopword case, since an empty tsquery emits a notice.

The vector leg#

384-dimension embeddings from bge-small-en-v1.5, L2-normalised, in a vector(384) column with an HNSW index. The dimension is hardcoded in the schema rather than left untyped, deliberately: the type check is what catches a model-mismatch bug at insert time, which is exactly the class of silent corruption the database should refuse.

hnsw.ef_search is set to 100 per query rather than left at pgvector's default of 40. The default is calibrated for an unfiltered index; once a selective tenant predicate discards most of the beam, a beam of 40 is too narrow and recall quietly drops on exactly the tenants with the least data.

Reciprocal rank fusion#

app/retrieval/fusion.py
score(d) = Σ_legs  weight_leg / (k + rank_leg(d))        # k = 60

The two legs produce scores on incompatible scales: ts_rank_cd is unbounded and depends on term frequencies in this corpus, cosine similarity is bounded and depends on the query. “0.83” means something entirely different in each. So fusion consumes only ranked id lists — it is scale-free, and there is no normalisation step to get wrong.

Why not weighted score fusion

The obvious approach, and worse than it looks. Min-max normalising each leg's returned window makes every leg's best result 1.0 and its worst 0.0, which destroys the information that a leg found nothing good. A leg whose top hit is garbage still contributes a confident 1.0. It also needs per-corpus weight tuning that has to be redone whenever the corpus changes. Z-scores need a distribution that is not available at query time, and keyword scores are nowhere near normal. Learning to rank is the right answer at scale and needs labelled data that 65 cases are not.

A broken leg returning nonsense can contribute at most weight/(k+1) per document, so it degrades the ranking gently instead of destroying it. The sort key is total — score, then leg count, then best rank, then chunk id — because two documents at rank 1 in different legs have identical scores, and without a deterministic final key their order depends on dictionary insertion. Evaluation metrics would then reshuffle between identical runs, which is a day spent chasing a phantom regression.

What fusion gives up: it is blind to margin. If the keyword leg's first result is an exact error-code hit and its second is unrelated, fusion treats that gap the same as a near-tie.

The cross-encoder#

bge-reranker-base scores query–chunk pairs jointly rather than comparing two independently-computed vectors, which is why it is more accurate and why it cannot be precomputed. It outputs a raw logit, roughly −11 to +11.

Both numbers are persisted: a sigmoid of the logit, which every gate and threshold reads, and the raw logit itself. A gate that never fires is debuggable only if you can see the distribution the model actually produced; keeping only the squashed value loses that, and keeping only the logit puts every downstream threshold on an unbounded model-specific scale. Sorting uses the raw logit, since a sigmoid is monotonic but saturates at the extremes.

Softmax over the candidate set was rejected: it makes each score relative to whatever else was retrieved alongside, so the same chunk scores differently depending on its company. A confidence gate needs an absolute signal.

Conditional reranking, and why it ships disabled

Reranking only when the top scores are ambiguous is the obvious latency optimisation. It is fully implemented, fully tested, and off by default — because always-rerank is the quality ceiling the evaluation measures the gate against. Shipping it on would make the gated arm the baseline and quietly delete the comparison. “Conditional reranking saves 280 ms and costs 1.2 points of recall@5” is a finding; “conditional reranking is on” is not.

The rule is margin = (s₁ − s₅) / s₁ over fused scores, because those depend only on ranks and k and therefore occupy the same numeric range for every query — the only thing that makes a relative threshold portable across queries.

Which reframes the open question from “what threshold?” to is the fusion margin a usable ambiguity signal at all? Fusion compresses scores hard by design — that is what k = 60 is for — so the dynamic range available to threshold on may simply be too narrow. If a sweep over 0.03–0.15 finds no threshold trading latency for acceptable recall, the alternative is gating on raw per-leg scores, at the cost of the cross-query comparability that made fused scores attractive in the first place. That is a real tension between two decisions in this system, and it should be resolved with data rather than argued.

Isolation, structurally#

Neither leg ever writes a full query. A leg supplies a projection, a predicate and an ordering, and the scope composes the real SQL with WHERE c.tenant_id = $1 AND c.is_current welded on unconditionally. $1 belongs to the scope; leg parameters start at $2.

The four layers that keep that honest →