Skip to content

Reference

Field notes

Seven bugs, each as expected → observed → root cause. The common thread is that none of them crashed: the full test suite was green the entire time, and every one of them would have shipped.
01

Keyword search returned zero rows, and nothing complained

What I expected

A realistic query like “webhook retry limit ERR_TIMEOUT_502” would match several chunks on the keyword leg and contribute them to fusion.

What happened

Zero rows. Rank fusion silently degenerated to vector-only, and “hybrid retrieval” became a claim in a README rather than something that happened. No error, no warning, and the arm you would have blamed is the one still working.

Why

Every convenient Postgres helper joins terms with AND. That is boolean retrieval: a document must contain every term or it does not match. BM25 sums a per-term contribution, so partial matches must still score. No single chunk held all six lexemes — the docs page explains retry behaviour and names the error code, while the changelog entry is the one that says “limit”.

What it changed

Fixing it by OR-ing everything then broke exact identifiers, because the parser splits ERR_TIMEOUT_502 into three lexemes and any chunk containing err matched. The right answer distinguishes & (concepts the user typed → OR) from <-> (one identifier held together → leave it alone). Three versions of one line, and only the second bug was findable by tests — the third needed looking at ranked output for a real query, which is what the playground is for.
02

The eval scored the wrong list, and made a working feature look useless

What I expected

The hybrid and hybrid+rerank arms would produce visibly different scorecards, since the second one spends 1.4 seconds per query on a cross-encoder.

What happened

Byte-identical scorecards. The reranker appeared to do nothing at all.

Why

The evaluation was scoring candidates — the pre-rerank list — because reranking returns a new list into results rather than re-sorting in place. Both arms had identical candidates, so both scored identically.

What it changed

The obvious conclusion from that scorecard would have been delete the reranker. Once fixed it was worth +12% MRR overall and +33% on normal questions. A measurement that confidently reports “no difference” is not neutral — it actively argues for removing the thing it failed to measure.
03

A test that guards against fake passes, which itself passed fakely

What I expected

The tenant-isolation test plants a secret in tenant B and asserts tenant A never sees it, with a control proving the secret is findable without the filter — so the test cannot pass on an empty index.

What happened

Green, and proving nothing.

Why

The control counted text-search matches across the whole chunks table. The real two-tenant corpus satisfied it easily, while the synthetic test tenants matched nothing at all. The control vouched for a corpus that was not the one under test.

What it changed

A control must be scoped as tightly as the thing it vouches for. It now asserts that both test tenants are reachable by the query. This is the most uncomfortable entry here, because the control existed specifically to prevent a vacuous pass, and vacuously passed.
04

A threshold nobody measured was wrong by 4x, in the direction that hid the feature

What I expected

A conditional-rerank margin of 0.30 would skip the cross-encoder on unambiguous queries and save meaningful latency.

What happened

It never fired. Real margins came out at 0.055–0.076, and the arithmetic caps the achievable value at 0.062.

Why

When every top-5 candidate is found by both legs — typical on this corpus — the fused scores are 2/(k+1) … 2/(k+5), a shape with a hard ceiling. So the margin was never measuring “how confident is the top result”. It measures “was the top five unanimous”, which is nearly binary.

What it changed

The number sounded defensible and sat in a config file with a comment explaining its reasoning. That is the whole problem: a value that looks tuned and does nothing is worse than an obvious placeholder, because it silently makes a comparison show no difference and invites the conclusion that the feature does not matter. Moved to 0.10, where it still never fires — so it is documented as implemented and unproven rather than as a working optimisation.
05

The same mistake, three times, in three places

What I expected

Three separate scorecards, each comparing two things.

What happened

Three confident numbers, none of which compared what it claimed to. The scorecard averaged three retrieval strategies into one row. The evaluation compared pre- and post-rerank lists as if they were the same list. The chunking experiment scored 8 cases in one arm and 41 in the other, and printed a delta anyway.

Why

All three are the same failure: a comparison that was not comparing the same things, presented as a confident number. Each one individually looked like an isolated slip.

What it changed

The fix was not more care — care had already been applied three times. It was a guard at each comparison point that checks validity before formatting: same arm, same list, same case count, or refuse to print a delta. When the same mistake appears three times, it is a missing mechanism, not three lapses.
06

A semantic cache is the most dangerous component in a RAG system

What I expected

A 0.95 similarity threshold is conservative enough that a semantic cache hit means essentially the same question.

What happened

ERR_TIMEOUT_502 and ERR_TIMEOUT_504 embed above 0.95 — they genuinely mean nearly the same thing — and have opposite correct answers.

Why

This is precisely the weakness hybrid retrieval exists to fix. Except the semantic cache is pure vector similarity: no keyword leg, no reranker, no confidence gate. It inherits the weakness with none of the mitigations that make it tolerable in retrieval.

What it changed

Identifier-bearing queries now skip the semantic cache entirely, on write as well as read, so weakening the read-side guard later cannot detonate a stored landmine. And abstentions are never cached: “I don't know” is a fact about the corpus at one moment, and caching it makes the system refuse documentation it now has.
07

Instrumentation that lies precisely when the feature works

What I expected

Replaying the cached answer's original cost onto a cache hit keeps the accounting honest.

What happened

Cost-per-query would rise as the cache hit rate improved. The dashboard would show the system getting more expensive exactly as it got cheaper.

Why

A cache hit is free. Attributing the original request's spend to it describes a request that never happened — and it does so on the one metric caching exists to move.

What it changed

A cache hit now records zero tokens and zero cost. Provider and model are kept, because “which model wrote this cached answer?” is a real debugging question. Instrumentation that inverts its signal is worse than none, because it is believed.

The pattern across all seven#

Five of these are the same species: something reported a number, and the number was about the wrong thing. A keyword leg reporting no matches because it was asking a boolean question. An evaluation reporting no difference because it was scoring the list before the change. A control reporting “findable” about a different corpus. A threshold reporting a decision it could never make. A cost metric reporting spend on a request that never ran.

None of them raised. All of them were confident. The defence that actually worked was not more tests — it was looking at real output for a real query, and building a harness whose job is to compare two things and refuse to print a number when they are not comparable.

What is still wrong, or unmeasured, or open →