Documentation · Overview
A support assistant that would rather say nothing than say something wrong.
Fishack answers questions from one customer's own documentation and refuses when that corpus cannot support an answer. Every claim carries a citation, every citation is checked after the answer is written, and a confidence gate escalates to a human instead of guessing.
- multi-tenant
- hybrid retrieval
- cited and validated
- confidence-gated
- no RAG framework
- 369 tests + 23 integration
- 65-case eval harness
The failure this was built against#
A support answer that sounds right is worse than no answer, because nobody checks it. The failure mode of a documentation assistant is not silence — it is a fluent paragraph assembled from two pages that contradict each other, with a version number silently dropped and no way for the reader to tell.
Fishack is built for a fictional B2B analytics and billing SaaS called Flowlytics, whose corpus contains the conflicts a real one does: product docs that disagree with a changelog, error codes one digit apart with opposite fixes, and support tickets whose question and resolution live in different places. Every defence in the system exists because one of those cases defeated the version before it.
The promise, and where you can see it#
The tagline is not marketing copy, it is a specification. Each clause names a property, and each property has a UI element whose only job is to make that property observable — and, more importantly, falsifiable.
| Clause | What proves it | How it could fail visibly |
|---|---|---|
| Every claim cited | Inline [n] markers in the answer, each one a button that opens the exact chunk it points at. | A marker pointing at a source the backend never offered renders as a fabrication, in rose, rather than as a working link. |
| Every citation verified | A per-claim verdict inside each source — supports, with the similarity, or weak match. | Validation runs after generation on every answer, and its failures are surfaced rather than suppressed. |
| Confidence-gated | A pill showing the score, the threshold, and which of the two scales the score is on. | The two scales differ by roughly 30x, so a bare number would be a lie. 0.02 is healthy on one and catastrophic on the other. |
| Escalates instead of guessing | An amber banner with the reason in plain language and a ticket id. | The ticket carries the conversation and the top ten sources with their scores, so a human does not repeat the search that just failed. |
| Tenant-isolated | A tenant switcher, and an answer that visibly declines to cross. | Switching clears the conversation, and says so before the click — carrying history across would feed one tenant's answers into another's prompt. |
How it fits together#
Everything on that diagram is written out. There is no LangChain, no LlamaIndex, and no vendor SDK anywhere in the system: retrieval, rank fusion, reranking, prompt assembly, citation validation, caching and evaluation are all code in this repository, and every LLM provider is raw httpx. That is the point of the project — the interesting decisions are precisely the ones a framework would have made for you, and they are the ones worth being able to defend.
What it deliberately is not#
- A general chatbot
- An out-of-corpus question gets an abstention and zero LLM calls — the confidence gate sits before generation, so nothing is spent on a question the corpus cannot answer.
- A framework demonstration
- No RAG framework is used. Every ranking function, threshold and prompt is in the repository, and each one has a comment saying how its value was chosen.
- A benchmark claim
- Every number here is measured on one 65-case golden set over one AI-written corpus, and each one says so. The limitations page states what that does and does not support.
- Production-tuned
- The cross-encoder costs more than the entire latency budget on this hardware, so it ships behind a flag. The measurement that justifies dropping it is a stronger artefact than shipping it on would have been.
What the evaluation found#
The harness lives in fishnet/ and is the most useful thing this project produced, because it contradicted its own design document. Three findings are worth arriving with.
Hybrid retrieval lost to vector-only on this corpus
| Arm | recall@5 | MRR | mean latency |
|---|---|---|---|
| BM25 only | 0.747 | 0.677 | 10 ms |
| Vector only | 0.938 | 0.820 | 40 ms |
| BM25 + vector (RRF) | 0.910 | 0.768 | 50 ms |
| BM25 + vector + reranker | 0.895 | 0.863 | 1,700 ms |
| Vector + reranker | 0.915 | 0.893 | 3,300 ms |
The sharpest single case: on “how long until my events show up in the dashboard?” the correct chunk sat at rank 7 under vector search and rank 20 after fusion. Only the top 8 candidates reach the reranker, so blending pushed the right answer out of the reranker's reach. The vector arm recovered it; the hybrid arm could not.
Per-source chunking beat fixed windows, decisively
The same corpus was ingested twice — once with the three per-source chunkers, once with fixed 1,600-character windows — under shadow tenants and scored identically. Naive chunking loses 45% of multi-turn answers entirely: not ranked lower, absent from the top 20. It cuts a ticket's question away from its resolution and strips the heading context that tells a chunk what it is about.
The cross-encoder costs more than the whole latency budget
Mean 1.7–3.3 seconds on a 12-thread CPU, against a target of P95 under three seconds for the entire request. Buying +0.073 MRR for +3.3 seconds is a real product decision, and the numbers to make it are now on the table rather than in a hunch. On a GPU or a hosted reranker that is roughly 200 ms and an obvious yes.
All measured results, with their provenance and caveats →