Evaluation
The eval harness
Why an eval harness is the deliverable#
Without one, every claim about a RAG system is an anecdote. “It seems better” after a chunker change is indistinguishable from the four queries you happened to try. More importantly, a system that can only confirm its own design is not being measured — and this harness earned its place by contradicting the design document that specified it, twice.
The golden set#
65 cases, derived from the corpus specification and then hand-edited. Six types, each of which exercises a different component.
| Case type | Cases | What it tests |
|---|---|---|
normal | 16 | Baseline — a direct question answered from one source |
out_of_scope | 16 | The confidence gate and the abstention path. MUST abstain |
exact_identifier | 9 | The keyword leg's entire justification. If these fail, keyword search is worthless |
multi_turn | 8 | Whether query rewriting works |
stale_conflict | 8 | Versioning and the conflict rule in the prompt |
cross_tenant | 8 | Tenant isolation. MUST NOT leak, and must also abstain |
What a case carries
class GoldenCase(BaseModel):
case_id: str # "GC-0041" — stable, human-assigned
case_type: CaseType
tenant_id: str
query: str
history: list[dict[str, str]] # prior turns, for multi_turn cases
expected_sources: list[SourceLocator] # NOT chunk ids — see below
reference_answer: str # prose: "a good answer looks like this"
must_contain: list[str] # ["5"] — literals that MUST appear
forbidden_text: str | None # cross_tenant: this string must never appear
notes: strTwo fields are each the result of one decision. reference_answer is prose rather than an exact string, because exact-match scoring on generated text measures phrasing, not correctness — “The retry limit is 5” and “Retries cap at 5” are both right. And must_contain is deterministic, which is the cheap complement to a noisy judge: it catches a fluent answer that quietly omits the actual figure. In a stale-data case, must_contain: ["5"] is the point — plausibly discussing retries is not enough.
Stable locators are the harness's most important idea
Ground truth is a SourceLocator — {source_type, slug, heading} for docs, an entry id for the changelog, a ticket id for tickets — resolved to chunk ids at the start of every run.
| Ground truth as… | Consequence |
|---|---|
| Chunk UUIDs | Simplest to compute, and wrong in a way that surfaces weeks later. Re-ingesting invalidates the whole golden set silently: every case points at rows that no longer exist, recall drops to zero, and it presents as a catastrophic retrieval regression with no cause |
| Content hashes | Stable across re-ingest only while the chunker produces byte-identical text. Breaks the moment chunk boundaries move — which is what the chunking experiment does |
| Source locators | Expresses ground truth in terms of the SOURCE, which both chunking arms share. Without this the chunking experiment is not awkward, it is impossible |
Resolution runs before any retrieval or generation, so a stale golden set fails in seconds rather than after twenty minutes of LLM calls. Unresolved locators are reported loudly and never dropped: a locator matching nothing means the golden set and the corpus have diverged, and without a signal that looks like a recall regression.
Four metrics, each with its arithmetic in the docstring#
| Metric | The question it answers | The decision inside it |
|---|---|---|
| recall@k | What fraction of the chunks that should have been retrieved made the top k? | Returns 1.0 when nothing was expected. An out-of-scope case had nothing to find, so nothing was missed — returning 0.0 would drag the aggregate down for the behaviour we WANT |
| precision@k | How much of the top k was actually relevant? | Divides by min(k, retrieved), not k — dividing by k punishes a system for returning three excellent results when the corpus only holds three. That is a penalty for a small corpus, not for being wrong |
| MRR | How high did the FIRST correct chunk land? | Two systems can share a recall@5 while one puts the answer first and the other fifth. Only MRR sees that, and language models are position-sensitive, so the difference is real |
| hit@k | Did any correct chunk make the top k? Yes or no | Coarser than recall and sometimes the more honest question. It became the chunking experiment's headline metric because it is immune to the measurement artefact described there |
The LLM judge#
Faithfulness, citation accuracy and answer relevance cannot be computed arithmetically, so they are scored by a model. Three things about how make the scores worth reading.
| Choice | Why |
|---|---|
| A different, larger model | A model judging its own output shares its blind spots. If the generator misreads a chunk, the same model asked 'is this faithful?' tends to misread it identically and say yes. The failures are correlated, so it measures agreement with itself rather than correctness |
| A separate provider chain, reversed | Sharing the chain means a rate-limited judge silently falls back onto the exact model it is grading — the correlated-failure problem reintroduced through the back door, invisibly |
| An explicit rubric with named criteria and anchors | 'Rate this 1-10' produces numbers that drift between runs and cannot be reasoned about. Named criteria with stated anchor points are reproducible enough to diff, and inspectable — when a score looks wrong you can read the rubric and see whether the judge or the rubric was at fault |
| The reference answer goes LAST in the prompt | Placed early it gets anchored on, and the judge scores similarity-to-reference rather than faithfulness-to-context — rewarding paraphrase over grounding |
1. FAITHFULNESS — is every factual claim in ANSWER supported by CONTEXT?
1.0 every claim traceable to the context
0.5 mostly supported, at least one claim goes beyond it
0.0 claims the context does not support at all
Judge against CONTEXT only, never against your own knowledge.
2. CITATION_ACCURACY — does each [n] marker point at a source that actually
supports the claim it is attached to? A claim with no marker counts against.
3. ANSWER_RELEVANCE — does ANSWER address the QUESTION?
An abstention scores 1.0 when the context genuinely lacks the answer, 0.0 when
it does not.
RULE: if sources conflict and ANSWER follows the newer one AND states that the
older one disagrees, that is full marks on faithfulness. It is the system's
designed behaviour, not a contradiction.Without that last rule the judge scores the system's best feature as a bug.
Hard assertions#
Fundamentally different from every other number here: binary, and with no tolerance band.
| Assertion | What it catches |
|---|---|
must_abstain | A confident, fluent, well-cited answer to a question the corpus cannot answer. This is the single failure the whole system exists to prevent, and the only check that measures what the gate, the closed-book prompt and the few-shot abstention examples are jointly trying to produce |
no_cross_tenant_leak | Two independent checks, because there are two ways to fail. A chunk-level check catches a retrieval leak; a text-level check catches content that reached the answer by some OTHER route — a wrongly-namespaced cache, a prompt-assembly bug, history carried across tenants. The chunk check cannot see those at all |
no_fabricated_citations | A [7] marker when only five sources were offered. Unambiguous — no model, no threshold — which is what makes it assertable rather than merely measurable |
must_contain | Required literals. The cheap deterministic complement to the judge's biggest weakness: a fluent answer that quietly omits the actual figure |
The CI gate#
| Quality metrics | Hard assertions | |
|---|---|---|
| What | recall@5, recall@20, MRR, faithfulness, citation accuracy | must-abstain, no cross-tenant leak, no fabricated citations |
| Shape | Continuous, noisy, they trade against each other | Binary |
| Tolerance | A >5% relative drop against the committed baseline fails | Zero — any failure fails the build |
There is no acceptable rate of cross-tenant leakage, and a percentage band on a security check is exactly how a bug gets absorbed by a quality budget. So the two kinds live in separate modules and aggregate separately, with the assertions at the top of the scorecard.
Three rules that matter more than the number
- Improvements never fail. Obvious until you write
abs(delta) > toleranceand start failing builds for getting better. - No baseline is a pass. A fresh clone must not block CI, or nobody can ever create the first baseline.
- Generation metrics are skipped when the judge ran on under 25% of cases. Comparing a score over eight cases against a baseline over sixty is not a comparison, and on a free tier quota-limited runs are common enough that failing on them would train people to ignore the gate — which costs more than the regressions it would catch.
The baseline is committed rather than taken from the previous run. Committing is a deliberate act with a reviewable diff: someone looked at a scorecard and decided it represented intended behaviour. Auto-updating from the last run lets quality erode one tolerated 4% drop at a time, each individually acceptable, none ever noticed.
Running it#
make eval-retrieval # five arms, no LLM calls, under a minute
make eval # full harness including the judge, vs the baseline
make chunking-experiment # naive vs per-source chunking, on shadow tenants
make tune # sweep the confidence gate
python -m fishnet.run --sample 10 # a subset, for a fast loop
python -m fishnet.run --resume # per-case checkpointing, for free-tier quota--resume exists because a full run on a free tier can be interrupted by quota exhaustion halfway through, and restarting from zero wastes the calls already spent. Every report also carries a snapshot of the configuration that produced it, so a scorecard can be read six weeks later without guessing which thresholds were in force.