Skip to content

Evaluation

Measured results

Every number this project claims, the command that reproduces it, and the caveat that bounds it. Two of these contradict the design document that specified the system, which is the most useful thing the harness produced.

Golden set

65

cases across six types and two tenants

Corpus

312

current chunks, 156 per tenant

Retrieval p50

29 ms

retrieval only, before any reranking

Test suite

369 + 23

unit plus integration

Retrieval strategies#

Five arms over the same 65 cases and the same corpus. Retrieval only, so there are no LLM calls and the run is fully reproducible on any machine with the repository checked out.

Armrecall@5recall@20MRRmean latency
BM25 only0.7470.9280.67710 ms
Vector onlybest recall@5 of any arm0.9381.0000.82040 ms
BM25 + vector (RRF)0.9100.9820.76850 ms
BM25 + vector + reranker0.8950.9820.8631,700 ms
Vector + rerankerbest MRR of any arm0.9151.0000.8933,300 ms
make eval-retrieval · macro-averaged over cases · both tenants.

Where the average is hiding the finding

The overall number conceals two opposite effects that happen to cancel. Broken out by case type, MRR:

Question typeBM25vectorhybrid
Exact identifiers (ERR_TIMEOUT_502)the keyword leg earns its place here, exactly as designed0.5420.6520.726
Multi-turn follow-upsand destroys the result here0.1590.6250.306

The keyword leg helps exactly where the design predicted and is near-useless on short, pronoun-heavy follow-ups. Equal-weight fusion then lets the harm win the average. Weighting the vector leg at 1.0 and the keyword leg at 0.5 is the obvious next experiment; it has not been run, and it is recorded as open rather than assumed.

The sharpest single case

On “how long until my events show up in the dashboard?” the correct chunk sat at rank 7 under vector search and rank 20 after fusion. Only the top eight candidates reach the reranker, so blending pushed the right answer out of the reranker's reach entirely. The vector arm recovered it; the hybrid arm could not.

That is a concrete mechanism rather than a statistical wobble, and it is the kind of finding that only exists because the harness reports per-case detail rather than a headline number.

Chunking: per-source versus fixed windows#

The same corpus ingested twice — once with the three per-source chunkers, once with fixed 1,600-character windows — under shadow tenants and scored identically.

Question typeMetricnaiveper-source
Overallrecall@50.5910.858
recall@200.8630.972
Exact identifiersrecall@50.6671.000
MRR0.4500.726
Multi-turnrecall@200.5501.000
make chunking-experiment · shadow tenants, so both arms are scored by the same golden set through source locators.

Naive chunking loses 45% of multi-turn answers entirely — not ranked lower, absent from the top twenty. It cuts a ticket's question away from its resolution and strips the heading context that tells a chunk what it is about. This is the largest single effect measured anywhere in the project, and it is on the ingestion side rather than the retrieval side.

Latency and cost#

retrieval only+ rerankerfull answer
p5029 ms1,548 msmeasured per run
p9571 ms5,290 msmeasured per run
12-thread CPU, no GPU. Full-answer latency depends on which provider served the request, so it is reported live on the operations dashboard rather than fixed here.

Cost is virtual, and says so everywhere

Real spend is $0.00 — every provider is on a free tier. So cost is tracked as what the same token usage would cost at paid-API list prices, priced per token from the model that actually served each request, against a price table snapshotted from public price pages.

The methodology label stays attached to every figure on the dashboard. Detached from it, “cost per query” reads as a bill. The figure exists to make the cost of a design choice visible before it is one.

What the harness caught that testing did not#

This list is the harness's return on investment, and the common thread is that the full test suite was green for every one of them.

FoundImpact
The eval was scoring the pre-rerank listMade a working feature look worthless — the obvious conclusion would have been to delete the reranker. Once fixed it was worth +12% MRR overall and +33% on normal questions
Hybrid retrieval losing to vector-onlyContradicted the design document. Reopened rank-fusion weighting as a live question
The conditional-rerank threshold was wrong by 4xA config value that looked tuned and could never fire
A leakage-test control passing vacuouslyThe security test was green and proving nothing
Three separate comparisons of non-comparable thingsAveraged strategies into one row; compared two different lists; scored 8 cases against 41. Each printed a confident number
cross_tenant cases were not required to abstainEight spurious warnings on the first run, and a weaker assertion than it should have been

Reproducing all of this#

ResultCommandNeeds a key?
Retrieval strategies tablemake eval-retrievalNo
Chunking experimentmake chunking-experimentNo
Latency percentilesmake eval-retrievalNo
Faithfulness and citation accuracymake evalYes — the judge is an LLM
Confidence threshold sweepmake tuneNo
Cost and live latencythe /admin dashboardYes, to generate traffic
Everything reproducible without a key is reproducible from a fresh clone, because the corpus and its generation cache are committed.

What these numbers do not support, stated before someone quotes them →