Skip to content

Documentation · Overview

A support assistant that would rather say nothing than say something wrong.

Fishack answers questions from one customer's own documentation and refuses when that corpus cannot support an answer. Every claim carries a citation, every citation is checked after the answer is written, and a confidence gate escalates to a human instead of guessing.

  • multi-tenant
  • hybrid retrieval
  • cited and validated
  • confidence-gated
  • no RAG framework
  • 369 tests + 23 integration
  • 65-case eval harness

The failure this was built against#

A support answer that sounds right is worse than no answer, because nobody checks it. The failure mode of a documentation assistant is not silence — it is a fluent paragraph assembled from two pages that contradict each other, with a version number silently dropped and no way for the reader to tell.

Fishack is built for a fictional B2B analytics and billing SaaS called Flowlytics, whose corpus contains the conflicts a real one does: product docs that disagree with a changelog, error codes one digit apart with opposite fixes, and support tickets whose question and resolution live in different places. Every defence in the system exists because one of those cases defeated the version before it.

The promise, and where you can see it#

The tagline is not marketing copy, it is a specification. Each clause names a property, and each property has a UI element whose only job is to make that property observable — and, more importantly, falsifiable.

ClauseWhat proves itHow it could fail visibly
Every claim citedInline [n] markers in the answer, each one a button that opens the exact chunk it points at.A marker pointing at a source the backend never offered renders as a fabrication, in rose, rather than as a working link.
Every citation verifiedA per-claim verdict inside each source — supports, with the similarity, or weak match.Validation runs after generation on every answer, and its failures are surfaced rather than suppressed.
Confidence-gatedA pill showing the score, the threshold, and which of the two scales the score is on.The two scales differ by roughly 30x, so a bare number would be a lie. 0.02 is healthy on one and catastrophic on the other.
Escalates instead of guessingAn amber banner with the reason in plain language and a ticket id.The ticket carries the conversation and the top ten sources with their scores, so a human does not repeat the search that just failed.
Tenant-isolatedA tenant switcher, and an answer that visibly declines to cross.Switching clears the conversation, and says so before the click — carrying history across would feed one tenant's answers into another's prompt.
Every row is a behaviour in the running product, not an aspiration. The guided tour walks all five in about three minutes.

How it fits together#

Fishack system architectureFour bands. The request plane holds the chat UI and the operations dashboard, both served by Next.js. Below it, the FastAPI query pipeline runs seven stages in order: rewrite, cache, retrieve, rerank, confidence gate, generate, validate. Below that, the data and model plane holds Postgres with pgvector, Redis, the local embedding and reranker models, and the four-provider LLM fallback chain. Underneath, two offline systems feed the pipeline: the ingestion pipeline that loads, chunks, embeds and versions the corpus, and the fishnet evaluation harness that scores retrieval and generation against a 65-case golden set and gates CI.Request plane · Next.js App RouterChat UI/try · streaming · sources panelOperations dashboard/admin · one row per requestSame-origin proxy — no CORS anywhere in the systemnext.config.mjs rewrites /api/* → FastAPI (ADR-027)Query path · FastAPI · app/generation/pipeline.py1. Rewritefollow-up →standalone2. Cacheexact, thensemantic3. RetrieveBM25 + vector→ RRF4. Rerankcross-encodertop 8 → 55. Gateconfidencethreshold6. Generateclosed-book,cited7. Validateper claim,post-hoccache hit → answer in ~10 ms, zero LLM callsbelow threshold → abstainData and model planePostgres 16 + pgvectorchunks · documents · traces · escalations · feedbackRedisanswer cache · quota countersLocal models (CPU)bge-small · bge-reranker-baseLLM fallback chainGroq → Gemini → OpenRouter → OllamaOffline — feeds the system, never serves a requestIngestion pipeline3 loaders → 3 chunking strategies → bge-small (cached)→ transactional insert → second passsupersede a doc · tag chunks a newer entry contestsidempotent by content hash · ~312 chunks, two tenantsfishnet — the evaluation harness65-case golden set → stable locators → recall@k · MRR→ LLM judge on a separate model and provider chain→ scorecard → CI gate against a committed baseline5% tolerance on quality · zero on correctness
Four bands: the request plane, the query pipeline, the data and model plane, and the two offline systems that feed it. Ingestion and the eval harness never serve a request — they produce the corpus the pipeline reads and the scorecard that says whether it got worse.

Everything on that diagram is written out. There is no LangChain, no LlamaIndex, and no vendor SDK anywhere in the system: retrieval, rank fusion, reranking, prompt assembly, citation validation, caching and evaluation are all code in this repository, and every LLM provider is raw httpx. That is the point of the project — the interesting decisions are precisely the ones a framework would have made for you, and they are the ones worth being able to defend.

What it deliberately is not#

A general chatbot
An out-of-corpus question gets an abstention and zero LLM calls — the confidence gate sits before generation, so nothing is spent on a question the corpus cannot answer.
A framework demonstration
No RAG framework is used. Every ranking function, threshold and prompt is in the repository, and each one has a comment saying how its value was chosen.
A benchmark claim
Every number here is measured on one 65-case golden set over one AI-written corpus, and each one says so. The limitations page states what that does and does not support.
Production-tuned
The cross-encoder costs more than the entire latency budget on this hardware, so it ships behind a flag. The measurement that justifies dropping it is a stronger artefact than shipping it on would have been.

What the evaluation found#

The harness lives in fishnet/ and is the most useful thing this project produced, because it contradicted its own design document. Three findings are worth arriving with.

Hybrid retrieval lost to vector-only on this corpus

Armrecall@5MRRmean latency
BM25 only0.7470.67710 ms
Vector only0.9380.82040 ms
BM25 + vector (RRF)0.9100.76850 ms
BM25 + vector + reranker0.8950.8631,700 ms
Vector + reranker0.9150.8933,300 ms
make eval-retrieval · 65 cases across both tenants · retrieval only, so no LLM calls and fully reproducible.

The sharpest single case: on “how long until my events show up in the dashboard?” the correct chunk sat at rank 7 under vector search and rank 20 after fusion. Only the top 8 candidates reach the reranker, so blending pushed the right answer out of the reranker's reach. The vector arm recovered it; the hybrid arm could not.

Per-source chunking beat fixed windows, decisively

The same corpus was ingested twice — once with the three per-source chunkers, once with fixed 1,600-character windows — under shadow tenants and scored identically. Naive chunking loses 45% of multi-turn answers entirely: not ranked lower, absent from the top 20. It cuts a ticket's question away from its resolution and strips the heading context that tells a chunk what it is about.

The cross-encoder costs more than the whole latency budget

Mean 1.7–3.3 seconds on a 12-thread CPU, against a target of P95 under three seconds for the entire request. Buying +0.073 MRR for +3.3 seconds is a real product decision, and the numbers to make it are now on the table rather than in a hunch. On a GPU or a hosted reranker that is roughly 200 ms and an obvious yes.

All measured results, with their provenance and caveats →

Where to go next#