Skip to content

Architecture

System architecture

Four planes — request, query, data, and the offline systems that feed them. Every component is written out rather than assembled from a framework, and every tuning knob has a recorded provenance saying whether its value was measured or guessed.
Fishack system architectureFour bands. The request plane holds the chat UI and the operations dashboard, both served by Next.js. Below it, the FastAPI query pipeline runs seven stages in order: rewrite, cache, retrieve, rerank, confidence gate, generate, validate. Below that, the data and model plane holds Postgres with pgvector, Redis, the local embedding and reranker models, and the four-provider LLM fallback chain. Underneath, two offline systems feed the pipeline: the ingestion pipeline that loads, chunks, embeds and versions the corpus, and the fishnet evaluation harness that scores retrieval and generation against a 65-case golden set and gates CI.Request plane · Next.js App RouterChat UI/try · streaming · sources panelOperations dashboard/admin · one row per requestSame-origin proxy — no CORS anywhere in the systemnext.config.mjs rewrites /api/* → FastAPI (ADR-027)Query path · FastAPI · app/generation/pipeline.py1. Rewritefollow-up →standalone2. Cacheexact, thensemantic3. RetrieveBM25 + vector→ RRF4. Rerankcross-encodertop 8 → 55. Gateconfidencethreshold6. Generateclosed-book,cited7. Validateper claim,post-hoccache hit → answer in ~10 ms, zero LLM callsbelow threshold → abstainData and model planePostgres 16 + pgvectorchunks · documents · traces · escalations · feedbackRedisanswer cache · quota countersLocal models (CPU)bge-small · bge-reranker-baseLLM fallback chainGroq → Gemini → OpenRouter → OllamaOffline — feeds the system, never serves a requestIngestion pipeline3 loaders → 3 chunking strategies → bge-small (cached)→ transactional insert → second passsupersede a doc · tag chunks a newer entry contestsidempotent by content hash · ~312 chunks, two tenantsfishnet — the evaluation harness65-case golden set → stable locators → recall@k · MRR→ LLM judge on a separate model and provider chain→ scorecard → CI gate against a committed baseline5% tolerance on quality · zero on correctness
The whole system. Ingestion and the evaluation harness sit underneath because they feed the pipeline rather than serving it: one produces the corpus the query path reads, the other produces the scorecard that says whether a change made it worse.

The shape of the system#

Fishack is one FastAPI application, one Next.js application, Postgres with pgvector, Redis, and two local transformer models. It is deliberately not a set of services: at this size a service boundary buys deployment independence nobody needs and costs a network hop, a serialisation format, and a failure mode on every call.

ComponentRoleNotable
Next.js App RouterDocumentation, the assistant, and the operations dashboardProxies every API call server-side, so the browser only ever talks to one origin and no CORS middleware exists anywhere
FastAPIThe query path, feedback, and admin statisticsLong-lived resources are built once in the lifespan context and read off app.state — no globals, so tests can build an app against fakes
Postgres 16 + pgvectorDocuments, chunks, embeddings, traces, escalations, feedbackOne datastore serves BOTH retrieval legs, so one tenant predicate covers both and index consistency is guaranteed by a generated column
RedisAnswer cache, cache reverse index, provider quota countersEvery call is wrapped: a Redis failure degrades to a cache miss rather than an error
Local models (CPU)bge-small embeddings, bge-reranker-base cross-encoderLoaded at startup, not on first request — so a missing model fails the deploy rather than the first customer
LLM fallback chainQuery rewriting, generation, and the eval judgeFour providers behind one interface, each raw httpx. Order is configuration; a quota-exhausted provider fails over without sleeping

Why there is no RAG framework

No LangChain, no LlamaIndex, no vendor SDKs. Retrieval, rank fusion, reranking, prompt assembly, citation validation, caching and evaluation are all written out, and every provider is raw httpx.

The reason is not purity. It is that the interesting decisions in a RAG system are exactly the ones a framework makes for you: how two incomparable score scales get merged, what a confidence threshold is measured against, whether an abstention is cacheable, which query the cache is keyed on. Each of those has a defensible answer here, and each has an ADR recording the alternatives that were rejected.

The request plane#

Every frontend call goes to a relative /api/… path, which the Next server rewrites onto FastAPI. Two consequences follow, and the second one is the real reason.

First, there is no CORS anywhere in the system — same-origin requests skip the mechanism entirely. That matters more than usual here because the chat response is a stream, and CORS combined with streaming is genuinely unpleasant to debug: a missing or wrong header produces a stream that simply stops, with no useful error in either the browser console or the server log.

Second, one environment variable covers two environments. In Docker the backend is api:8000 on the compose network; locally it is localhost:8000. Only API_ORIGIN changes, it is read on the Next server, and it never reaches the browser — which has no idea the backend exists as a separate thing.

Storage layout#

Five tables and one cache, with the versioning columns on documents and the tenant column denormalised onto chunks so that the isolation predicate never needs a join.

app/db/migrations/001_init.sql (shape)
tenants
  id                text primary key

documents
  id                uuid primary key
  tenant_id         text  references tenants
  source_type       text        -- docs | changelog | ticket
  source_path       text
  doc_version       text
  effective_date    date
  content_hash      text        -- dedup: re-ingesting identical content is a no-op
  is_current        boolean     -- superseded documents are archived, never deleted

chunks
  id                uuid primary key
  document_id       uuid  references documents
  tenant_id         text        -- DENORMALISED: the isolation predicate never joins
  chunk_index       int
  content           text        -- heading path is prepended into this (ADR-004)
  heading_path      text
  metadata          jsonb       -- doc_version, product_area, error_code, conflicts_with_entry…
  embedding         vector(384) -- HNSW index
  tsv               tsvector    -- GENERATED from content; GIN index
  is_current        boolean
  unique (document_id, chunk_index)

embedding_cache     -- sha256(model + text) -> vector. Makes re-ingestion nearly free.
traces              -- one row per request, written after the response
escalations         -- one row per abstention, with history and the top 10 chunks + scores
feedback            -- thumbs, keyed by trace_id

Why tenant_id is denormalised onto chunks

Because the isolation predicate is on the hot path of every single read, and a join is a thing a developer can forget to write. With tenant_id on the row, the scope that owns the FROM clause can weld the predicate on unconditionally, and the runtime tripwire that re-checks every returned row has something local to check against.

What a trace row carries, and why

One row per request, written once after the response: the action taken, the confidence and which scale it is on, the retrieved chunk ids, the citation report as JSONB with its grounding rate precomputed, provider and model, tokens, virtual cost, and per-stage latencies. Writing traces never raises — an observability failure must not become a user-facing one.

The score kind is on the row rather than inferred later because top_score = 0.02 is uninterpretable six weeks after the fact. It is a healthy fusion score and a catastrophic reranker score, and a trace that cannot be read is not observability.

Every tuning knob, and where its value came from#

The status column is the point of this table. A number that looks tuned and never was is worse than an obvious placeholder — one of them was wrong by 4x in the direction that made its own feature invisible, and it sat in configuration with a confident comment explaining its reasoning.

SettingValueProvenanceReasoning
retrieval_candidates_per_leg20designAsking each leg for more than we keep is nearly free and gives fusion room to promote a chunk one leg ranked 18th and the other 2nd
retrieval_fusion_top_k20designThe candidate set that recall@20 is computed over
rerank_input_top_k8measuredThe cross-encoder costs ~270 ms per pair on CPU. Eight instead of twenty is ~2.5x faster and retrieval is unaffected, because fusion still emits twenty
rerank_top_k5designMore context costs tokens, latency, and needle-in-haystack dilution
rrf_k60literatureCormack et al. 2009. Damps rank 1 against rank 2 to under 2%, so one leg cannot steamroll the other
rrf_weight_bm25 / _vector1.0 / 1.0open questionWeights were avoided on principle, then evaluation measured BM25 actively hurting multi-turn recall. 0.5 is the obvious next experiment and has not been run
hnsw_ef_search100designpgvector's default of 40 is too narrow once a selective tenant filter discards most of the beam
conditional_rerank_enabledfalsedeliberateAlways-rerank is the quality ceiling the evaluation measures the gate against. Shipping the gate on would make the gated arm the baseline and delete the comparison
rerank_margin_threshold0.10measured, still not firingThe original 0.30 was wrong by 4x. Real margins are 0.055–0.076, and even 0.10 never fires — so conditional reranking is implemented and not yet demonstrated to do anything on this corpus
confidence_threshold_rerank0.45guessProvisional and labelled so. `make tune` sweeps it against the golden set
confidence_threshold_fused0.015guessSeparate from the above because the two scales differ by roughly 30x
semantic_cache_threshold0.95design + guardrailsSurvivable only because identifier-bearing queries skip the semantic cache entirely and abstentions are never cached
citation_similarity_threshold0.50deliberateLenient on purpose — it is catching UNRELATED citations, not grading paraphrase quality
cache_ttl_seconds3600designShorter than the freshness requirement. It is the backstop for bugs in active invalidation, not the primary mechanism
All of these live in app/config.py, which is the single place any constant, threshold or model name is allowed to exist. Model names are configuration rather than code because free-tier lineups change monthly.

The configuration reference lists every environment variable alongside these →

Build order#

The system was built in seven phases, each of which had to be demonstrably working before the next began. The ordering is not arbitrary: infrastructure and the LLM client came first so that everything after could assume a working provider chain; the evaluation harness came after the pipeline so it had something real to score, and immediately contradicted two of the pipeline's design assumptions.

PhaseWhat it addedDecisions
0 · FoundationsSchema and migrations, the four-provider LLM client with retries and failover, budget tracking, /healthADR-001, 002, 003, 005, 006
1 · IngestionThree loaders, three chunking strategies, cached embeddings, versioning and conflict taggingADR-004, 008, 009, 010
2 · RetrievalThe keyword and vector legs, rank fusion, the cross-encoder, and the isolation coreADR-011, 012, 013, 014
3 · GenerationQuery rewriting, the confidence gate, grounded generation, citation validation, escalations, tracesADR-015, 016, 017, 018
4 · EvaluationThe golden set, stable locators, the metrics, the LLM judge, the CI regression gateADR-019, 020, 021, 022
5 · Cache & feedbackExact and semantic caching, active invalidation, feedback triage, the admin statistics endpointADR-023, 024, 025, 026
6 · FrontendThe assistant, the operations dashboard, and this documentation siteADR-027, 028, 029