Reference
Field notes
Keyword search returned zero rows, and nothing complained
What I expected
What happened
Why
AND. That is boolean retrieval: a document must contain every term or it does not match. BM25 sums a per-term contribution, so partial matches must still score. No single chunk held all six lexemes — the docs page explains retry behaviour and names the error code, while the changelog entry is the one that says “limit”.What it changed
ERR_TIMEOUT_502 into three lexemes and any chunk containing err matched. The right answer distinguishes & (concepts the user typed → OR) from <-> (one identifier held together → leave it alone). Three versions of one line, and only the second bug was findable by tests — the third needed looking at ranked output for a real query, which is what the playground is for.The eval scored the wrong list, and made a working feature look useless
What I expected
hybrid and hybrid+rerank arms would produce visibly different scorecards, since the second one spends 1.4 seconds per query on a cross-encoder.What happened
Why
candidates — the pre-rerank list — because reranking returns a new list into results rather than re-sorting in place. Both arms had identical candidates, so both scored identically.What it changed
A test that guards against fake passes, which itself passed fakely
What I expected
What happened
Why
What it changed
A threshold nobody measured was wrong by 4x, in the direction that hid the feature
What I expected
What happened
Why
2/(k+1) … 2/(k+5), a shape with a hard ceiling. So the margin was never measuring “how confident is the top result”. It measures “was the top five unanimous”, which is nearly binary.What it changed
The same mistake, three times, in three places
What I expected
What happened
Why
What it changed
A semantic cache is the most dangerous component in a RAG system
What I expected
What happened
ERR_TIMEOUT_502 and ERR_TIMEOUT_504 embed above 0.95 — they genuinely mean nearly the same thing — and have opposite correct answers.Why
What it changed
Instrumentation that lies precisely when the feature works
What I expected
What happened
Why
What it changed
The pattern across all seven#
Five of these are the same species: something reported a number, and the number was about the wrong thing. A keyword leg reporting no matches because it was asking a boolean question. An evaluation reporting no difference because it was scoring the list before the change. A control reporting “findable” about a different corpus. A threshold reporting a decision it could never make. A cost metric reporting spend on a request that never ran.
None of them raised. All of them were confident. The defence that actually worked was not more tests — it was looking at real output for a real query, and building a harness whose job is to compare two things and refuse to print a number when they are not comparable.