Skip to content

Evaluation

The eval harness

65 cases across six types, ground truth expressed as source locators rather than chunk ids, four metrics whose arithmetic can be checked by hand, an LLM judge on a separate model and provider chain, and a CI gate with two tolerances — 5% on quality and zero on correctness.

Why an eval harness is the deliverable#

Without one, every claim about a RAG system is an anecdote. “It seems better” after a chunker change is indistinguishable from the four queries you happened to try. More importantly, a system that can only confirm its own design is not being measured — and this harness earned its place by contradicting the design document that specified it, twice.

The golden set#

65 cases, derived from the corpus specification and then hand-edited. Six types, each of which exercises a different component.

Case typeCasesWhat it tests
normal16Baseline — a direct question answered from one source
out_of_scope16The confidence gate and the abstention path. MUST abstain
exact_identifier9The keyword leg's entire justification. If these fail, keyword search is worthless
multi_turn8Whether query rewriting works
stale_conflict8Versioning and the conflict rule in the prompt
cross_tenant8Tenant isolation. MUST NOT leak, and must also abstain
Every metric is reported per case type as well as overall — see below for why that is not a presentation choice.

What a case carries

fishnet/models.py
class GoldenCase(BaseModel):
    case_id: str                    # "GC-0041" — stable, human-assigned
    case_type: CaseType
    tenant_id: str
    query: str
    history: list[dict[str, str]]   # prior turns, for multi_turn cases

    expected_sources: list[SourceLocator]   # NOT chunk ids — see below

    reference_answer: str           # prose: "a good answer looks like this"
    must_contain: list[str]         # ["5"] — literals that MUST appear
    forbidden_text: str | None      # cross_tenant: this string must never appear
    notes: str

Two fields are each the result of one decision. reference_answer is prose rather than an exact string, because exact-match scoring on generated text measures phrasing, not correctness — “The retry limit is 5” and “Retries cap at 5” are both right. And must_contain is deterministic, which is the cheap complement to a noisy judge: it catches a fluent answer that quietly omits the actual figure. In a stale-data case, must_contain: ["5"] is the point — plausibly discussing retries is not enough.

Stable locators are the harness's most important idea

Ground truth is a SourceLocator — {source_type, slug, heading} for docs, an entry id for the changelog, a ticket id for tickets — resolved to chunk ids at the start of every run.

Ground truth as…Consequence
Chunk UUIDsSimplest to compute, and wrong in a way that surfaces weeks later. Re-ingesting invalidates the whole golden set silently: every case points at rows that no longer exist, recall drops to zero, and it presents as a catastrophic retrieval regression with no cause
Content hashesStable across re-ingest only while the chunker produces byte-identical text. Breaks the moment chunk boundaries move — which is what the chunking experiment does
Source locatorsExpresses ground truth in terms of the SOURCE, which both chunking arms share. Without this the chunking experiment is not awkward, it is impossible

Resolution runs before any retrieval or generation, so a stale golden set fails in seconds rather than after twenty minutes of LLM calls. Unresolved locators are reported loudly and never dropped: a locator matching nothing means the golden set and the corpus have diverged, and without a signal that looks like a recall regression.

Four metrics, each with its arithmetic in the docstring#

MetricThe question it answersThe decision inside it
recall@kWhat fraction of the chunks that should have been retrieved made the top k?Returns 1.0 when nothing was expected. An out-of-scope case had nothing to find, so nothing was missed — returning 0.0 would drag the aggregate down for the behaviour we WANT
precision@kHow much of the top k was actually relevant?Divides by min(k, retrieved), not k — dividing by k punishes a system for returning three excellent results when the corpus only holds three. That is a penalty for a small corpus, not for being wrong
MRRHow high did the FIRST correct chunk land?Two systems can share a recall@5 while one puts the answer first and the other fifth. Only MRR sees that, and language models are position-sensitive, so the difference is real
hit@kDid any correct chunk make the top k? Yes or noCoarser than recall and sometimes the more honest question. It became the chunking experiment's headline metric because it is immune to the measurement artefact described there

The LLM judge#

Faithfulness, citation accuracy and answer relevance cannot be computed arithmetically, so they are scored by a model. Three things about how make the scores worth reading.

ChoiceWhy
A different, larger modelA model judging its own output shares its blind spots. If the generator misreads a chunk, the same model asked 'is this faithful?' tends to misread it identically and say yes. The failures are correlated, so it measures agreement with itself rather than correctness
A separate provider chain, reversedSharing the chain means a rate-limited judge silently falls back onto the exact model it is grading — the correlated-failure problem reintroduced through the back door, invisibly
An explicit rubric with named criteria and anchors'Rate this 1-10' produces numbers that drift between runs and cannot be reasoned about. Named criteria with stated anchor points are reproducible enough to diff, and inspectable — when a score looks wrong you can read the rubric and see whether the judge or the rubric was at fault
The reference answer goes LAST in the promptPlaced early it gets anchored on, and the judge scores similarity-to-reference rather than faithfulness-to-context — rewarding paraphrase over grounding
fishnet/judge.py — the rubric, and one rule specific to this system
1. FAITHFULNESS — is every factual claim in ANSWER supported by CONTEXT?
   1.0  every claim traceable to the context
   0.5  mostly supported, at least one claim goes beyond it
   0.0  claims the context does not support at all
   Judge against CONTEXT only, never against your own knowledge.

2. CITATION_ACCURACY — does each [n] marker point at a source that actually
   supports the claim it is attached to? A claim with no marker counts against.

3. ANSWER_RELEVANCE — does ANSWER address the QUESTION?
   An abstention scores 1.0 when the context genuinely lacks the answer, 0.0 when
   it does not.

RULE: if sources conflict and ANSWER follows the newer one AND states that the
older one disagrees, that is full marks on faithfulness. It is the system's
designed behaviour, not a contradiction.

Without that last rule the judge scores the system's best feature as a bug.

Hard assertions#

Fundamentally different from every other number here: binary, and with no tolerance band.

AssertionWhat it catches
must_abstainA confident, fluent, well-cited answer to a question the corpus cannot answer. This is the single failure the whole system exists to prevent, and the only check that measures what the gate, the closed-book prompt and the few-shot abstention examples are jointly trying to produce
no_cross_tenant_leakTwo independent checks, because there are two ways to fail. A chunk-level check catches a retrieval leak; a text-level check catches content that reached the answer by some OTHER route — a wrongly-namespaced cache, a prompt-assembly bug, history carried across tenants. The chunk check cannot see those at all
no_fabricated_citationsA [7] marker when only five sources were offered. Unambiguous — no model, no threshold — which is what makes it assertable rather than merely measurable
must_containRequired literals. The cheap deterministic complement to the judge's biggest weakness: a fluent answer that quietly omits the actual figure
Checks that do not apply to a case return nothing and are dropped — '3 of 3 passed' must never include skipped checks.

The CI gate#

Quality metricsHard assertions
Whatrecall@5, recall@20, MRR, faithfulness, citation accuracymust-abstain, no cross-tenant leak, no fabricated citations
ShapeContinuous, noisy, they trade against each otherBinary
ToleranceA >5% relative drop against the committed baseline failsZero — any failure fails the build

There is no acceptable rate of cross-tenant leakage, and a percentage band on a security check is exactly how a bug gets absorbed by a quality budget. So the two kinds live in separate modules and aggregate separately, with the assertions at the top of the scorecard.

Three rules that matter more than the number

  • Improvements never fail. Obvious until you write abs(delta) > tolerance and start failing builds for getting better.
  • No baseline is a pass. A fresh clone must not block CI, or nobody can ever create the first baseline.
  • Generation metrics are skipped when the judge ran on under 25% of cases. Comparing a score over eight cases against a baseline over sixty is not a comparison, and on a free tier quota-limited runs are common enough that failing on them would train people to ignore the gate — which costs more than the regressions it would catch.

The baseline is committed rather than taken from the previous run. Committing is a deliberate act with a reviewable diff: someone looked at a scorecard and decided it represented intended behaviour. Auto-updating from the last run lets quality erode one tolerated 4% drop at a time, each individually acceptable, none ever noticed.

Running it#

shell
make eval-retrieval          # five arms, no LLM calls, under a minute
make eval                    # full harness including the judge, vs the baseline
make chunking-experiment     # naive vs per-source chunking, on shadow tenants
make tune                    # sweep the confidence gate

python -m fishnet.run --sample 10        # a subset, for a fast loop
python -m fishnet.run --resume           # per-case checkpointing, for free-tier quota

--resume exists because a full run on a free tier can be interrupted by quota exhaustion halfway through, and restarting from zero wastes the calls already spent. Every report also carries a snapshot of the configuration that produced it, so a scorecard can be read six weeks later without guessing which thresholds were in force.

Every result the harness has produced →