Architecture
Retrieval and ranking
The keyword leg#
BM25-style keyword matching is provided by Postgres full-text search: a generated tsvector column on chunks with a GIN index, ranked with ts_rank_cd.
The alternative was a real BM25 engine. A Python BM25 library keeps its index in process memory, so it must be rebuilt on every restart and for every worker, and tenant filtering happens in Python after scoring — which is precisely the application-level isolation this system is built to avoid. Elasticsearch is real BM25 and a second datastore to keep consistent with Postgres, which is a large operational bill at this scale.
One datastore for both legs means one tenant predicate enforced in SQL for both, one backup story, and index consistency guaranteed by the generated column.
The query must OR its terms, and every convenient helper ANDs them
This shipped broken and is worth recording honestly, because the failure was silent in exactly the way this project is meant to guard against.
websearch_to_tsquery('english', 'webhook retry limit')
-> 'webhook' & 'retri' & 'limit'That is boolean retrieval: a document must contain every term or it does not match at all. BM25 does not work that way — it sums a per-term contribution over the terms a document does contain, so a document matching four of six query terms still scores, just lower. Ranked retrieval is inherently OR-ed; the ranking function decides what wins, not the matcher.
The consequence was severe and quiet. The realistic query "webhook retry limit ERR_TIMEOUT_502" returned zero rows, because the docs page explains retry behaviour and names the error code while the changelog entry is the one that says “limit”. No single chunk held all six lexemes. The keyword leg contributed nothing, fusion degenerated to vector-only, and “hybrid retrieval” became a claim in a README rather than something that happened.
Two operators, two meanings
The naive fix — OR everything — broke exact identifiers instead. Postgres' default parser treats _ as a separator, so ERR_TIMEOUT_502 becomes three independent lexemes and a flat OR matches any chunk containing err. On the real corpus, a query for one error code pulled a completely different error code's ticket into the top five.
The observation that resolves it is that websearch_to_tsquery already distinguishes the two cases:
websearch_to_tsquery('english', 'ERR_TIMEOUT_502')
-> 'err' <-> 'timeout' <-> '502' -- followed-by: adjacency IS the point| Operator | What it separates | Correct handling |
|---|---|---|
& | Concepts the user typed as separate ideas | Rewrite to | — this is ranked retrieval |
<-> | One identifier, or a quoted phrase, that the tokenizer happened to split | Leave alone — adjacency is the signal |
SELECT NULLIF(replace(websearch_to_tsquery('english', $2)::text, ' & ', ' | '), '')::tsquery
-- 'webhook' & 'retri' & 'limit' & 'err' <-> 'timeout' <-> '502'
-- becomes
-- 'webhook' | 'retri' | 'limit' | ('err' <-> 'timeout' <-> '502')Operator precedence does the rest: <-> binds tighter than &, which binds tighter than |, so the phrase group survives intact. A chunk mentioning only err no longer matches the identifier clause; a chunk about webhooks still matches on one concept. NULLIF handles the all-stopword case, since an empty tsquery emits a notice.
The vector leg#
384-dimension embeddings from bge-small-en-v1.5, L2-normalised, in a vector(384) column with an HNSW index. The dimension is hardcoded in the schema rather than left untyped, deliberately: the type check is what catches a model-mismatch bug at insert time, which is exactly the class of silent corruption the database should refuse.
hnsw.ef_search is set to 100 per query rather than left at pgvector's default of 40. The default is calibrated for an unfiltered index; once a selective tenant predicate discards most of the beam, a beam of 40 is too narrow and recall quietly drops on exactly the tenants with the least data.
Reciprocal rank fusion#
score(d) = Σ_legs weight_leg / (k + rank_leg(d)) # k = 60The two legs produce scores on incompatible scales: ts_rank_cd is unbounded and depends on term frequencies in this corpus, cosine similarity is bounded and depends on the query. “0.83” means something entirely different in each. So fusion consumes only ranked id lists — it is scale-free, and there is no normalisation step to get wrong.
Why not weighted score fusion
The obvious approach, and worse than it looks. Min-max normalising each leg's returned window makes every leg's best result 1.0 and its worst 0.0, which destroys the information that a leg found nothing good. A leg whose top hit is garbage still contributes a confident 1.0. It also needs per-corpus weight tuning that has to be redone whenever the corpus changes. Z-scores need a distribution that is not available at query time, and keyword scores are nowhere near normal. Learning to rank is the right answer at scale and needs labelled data that 65 cases are not.
A broken leg returning nonsense can contribute at most weight/(k+1) per document, so it degrades the ranking gently instead of destroying it. The sort key is total — score, then leg count, then best rank, then chunk id — because two documents at rank 1 in different legs have identical scores, and without a deterministic final key their order depends on dictionary insertion. Evaluation metrics would then reshuffle between identical runs, which is a day spent chasing a phantom regression.
What fusion gives up: it is blind to margin. If the keyword leg's first result is an exact error-code hit and its second is unrelated, fusion treats that gap the same as a near-tie.
The cross-encoder#
bge-reranker-base scores query–chunk pairs jointly rather than comparing two independently-computed vectors, which is why it is more accurate and why it cannot be precomputed. It outputs a raw logit, roughly −11 to +11.
Both numbers are persisted: a sigmoid of the logit, which every gate and threshold reads, and the raw logit itself. A gate that never fires is debuggable only if you can see the distribution the model actually produced; keeping only the squashed value loses that, and keeping only the logit puts every downstream threshold on an unbounded model-specific scale. Sorting uses the raw logit, since a sigmoid is monotonic but saturates at the extremes.
Softmax over the candidate set was rejected: it makes each score relative to whatever else was retrieved alongside, so the same chunk scores differently depending on its company. A confidence gate needs an absolute signal.
Conditional reranking, and why it ships disabled
Reranking only when the top scores are ambiguous is the obvious latency optimisation. It is fully implemented, fully tested, and off by default — because always-rerank is the quality ceiling the evaluation measures the gate against. Shipping it on would make the gated arm the baseline and quietly delete the comparison. “Conditional reranking saves 280 ms and costs 1.2 points of recall@5” is a finding; “conditional reranking is on” is not.
The rule is margin = (s₁ − s₅) / s₁ over fused scores, because those depend only on ranks and k and therefore occupy the same numeric range for every query — the only thing that makes a relative threshold portable across queries.
Which reframes the open question from “what threshold?” to is the fusion margin a usable ambiguity signal at all? Fusion compresses scores hard by design — that is what k = 60 is for — so the dynamic range available to threshold on may simply be too narrow. If a sweep over 0.03–0.15 finds no threshold trading latency for acceptable recall, the alternative is gating on raw per-leg scores, at the cost of the cross-query comparability that made fused scores attractive in the first place. That is a real tension between two decisions in this system, and it should be resolved with data rather than argued.
Isolation, structurally#
Neither leg ever writes a full query. A leg supplies a projection, a predicate and an ordering, and the scope composes the real SQL with WHERE c.tenant_id = $1 AND c.is_current welded on unconditionally. $1 belongs to the scope; leg parameters start at $2.