Evaluation
Measured results
Golden set
65
cases across six types and two tenants
Corpus
312
current chunks, 156 per tenant
Retrieval p50
29 ms
retrieval only, before any reranking
Test suite
369 + 23
unit plus integration
Retrieval strategies#
Five arms over the same 65 cases and the same corpus. Retrieval only, so there are no LLM calls and the run is fully reproducible on any machine with the repository checked out.
| Arm | recall@5 | recall@20 | MRR | mean latency |
|---|---|---|---|---|
| BM25 only | 0.747 | 0.928 | 0.677 | 10 ms |
| Vector onlybest recall@5 of any arm | 0.938 | 1.000 | 0.820 | 40 ms |
| BM25 + vector (RRF) | 0.910 | 0.982 | 0.768 | 50 ms |
| BM25 + vector + reranker | 0.895 | 0.982 | 0.863 | 1,700 ms |
| Vector + rerankerbest MRR of any arm | 0.915 | 1.000 | 0.893 | 3,300 ms |
Where the average is hiding the finding
The overall number conceals two opposite effects that happen to cancel. Broken out by case type, MRR:
| Question type | BM25 | vector | hybrid |
|---|---|---|---|
| Exact identifiers (ERR_TIMEOUT_502)the keyword leg earns its place here, exactly as designed | 0.542 | 0.652 | 0.726 |
| Multi-turn follow-upsand destroys the result here | 0.159 | 0.625 | 0.306 |
The keyword leg helps exactly where the design predicted and is near-useless on short, pronoun-heavy follow-ups. Equal-weight fusion then lets the harm win the average. Weighting the vector leg at 1.0 and the keyword leg at 0.5 is the obvious next experiment; it has not been run, and it is recorded as open rather than assumed.
The sharpest single case
On “how long until my events show up in the dashboard?” the correct chunk sat at rank 7 under vector search and rank 20 after fusion. Only the top eight candidates reach the reranker, so blending pushed the right answer out of the reranker's reach entirely. The vector arm recovered it; the hybrid arm could not.
That is a concrete mechanism rather than a statistical wobble, and it is the kind of finding that only exists because the harness reports per-case detail rather than a headline number.
Chunking: per-source versus fixed windows#
The same corpus ingested twice — once with the three per-source chunkers, once with fixed 1,600-character windows — under shadow tenants and scored identically.
| Question type | Metric | naive | per-source |
|---|---|---|---|
| Overall | recall@5 | 0.591 | 0.858 |
| recall@20 | 0.863 | 0.972 | |
| Exact identifiers | recall@5 | 0.667 | 1.000 |
| MRR | 0.450 | 0.726 | |
| Multi-turn | recall@20 | 0.550 | 1.000 |
Naive chunking loses 45% of multi-turn answers entirely — not ranked lower, absent from the top twenty. It cuts a ticket's question away from its resolution and strips the heading context that tells a chunk what it is about. This is the largest single effect measured anywhere in the project, and it is on the ingestion side rather than the retrieval side.
Latency and cost#
| retrieval only | + reranker | full answer | |
|---|---|---|---|
| p50 | 29 ms | 1,548 ms | measured per run |
| p95 | 71 ms | 5,290 ms | measured per run |
Cost is virtual, and says so everywhere
Real spend is $0.00 — every provider is on a free tier. So cost is tracked as what the same token usage would cost at paid-API list prices, priced per token from the model that actually served each request, against a price table snapshotted from public price pages.
The methodology label stays attached to every figure on the dashboard. Detached from it, “cost per query” reads as a bill. The figure exists to make the cost of a design choice visible before it is one.
What the harness caught that testing did not#
This list is the harness's return on investment, and the common thread is that the full test suite was green for every one of them.
| Found | Impact |
|---|---|
| The eval was scoring the pre-rerank list | Made a working feature look worthless — the obvious conclusion would have been to delete the reranker. Once fixed it was worth +12% MRR overall and +33% on normal questions |
| Hybrid retrieval losing to vector-only | Contradicted the design document. Reopened rank-fusion weighting as a live question |
| The conditional-rerank threshold was wrong by 4x | A config value that looked tuned and could never fire |
| A leakage-test control passing vacuously | The security test was green and proving nothing |
| Three separate comparisons of non-comparable things | Averaged strategies into one row; compared two different lists; scored 8 cases against 41. Each printed a confident number |
| cross_tenant cases were not required to abstain | Eight spurious warnings on the first run, and a weaker assertion than it should have been |
Reproducing all of this#
| Result | Command | Needs a key? |
|---|---|---|
| Retrieval strategies table | make eval-retrieval | No |
| Chunking experiment | make chunking-experiment | No |
| Latency percentiles | make eval-retrieval | No |
| Faithfulness and citation accuracy | make eval | Yes — the judge is an LLM |
| Confidence threshold sweep | make tune | No |
| Cost and live latency | the /admin dashboard | Yes, to generate traffic |
What these numbers do not support, stated before someone quotes them →