Architecture
System architecture
The shape of the system#
Fishack is one FastAPI application, one Next.js application, Postgres with pgvector, Redis, and two local transformer models. It is deliberately not a set of services: at this size a service boundary buys deployment independence nobody needs and costs a network hop, a serialisation format, and a failure mode on every call.
| Component | Role | Notable |
|---|---|---|
| Next.js App Router | Documentation, the assistant, and the operations dashboard | Proxies every API call server-side, so the browser only ever talks to one origin and no CORS middleware exists anywhere |
| FastAPI | The query path, feedback, and admin statistics | Long-lived resources are built once in the lifespan context and read off app.state — no globals, so tests can build an app against fakes |
| Postgres 16 + pgvector | Documents, chunks, embeddings, traces, escalations, feedback | One datastore serves BOTH retrieval legs, so one tenant predicate covers both and index consistency is guaranteed by a generated column |
| Redis | Answer cache, cache reverse index, provider quota counters | Every call is wrapped: a Redis failure degrades to a cache miss rather than an error |
| Local models (CPU) | bge-small embeddings, bge-reranker-base cross-encoder | Loaded at startup, not on first request — so a missing model fails the deploy rather than the first customer |
| LLM fallback chain | Query rewriting, generation, and the eval judge | Four providers behind one interface, each raw httpx. Order is configuration; a quota-exhausted provider fails over without sleeping |
Why there is no RAG framework
No LangChain, no LlamaIndex, no vendor SDKs. Retrieval, rank fusion, reranking, prompt assembly, citation validation, caching and evaluation are all written out, and every provider is raw httpx.
The reason is not purity. It is that the interesting decisions in a RAG system are exactly the ones a framework makes for you: how two incomparable score scales get merged, what a confidence threshold is measured against, whether an abstention is cacheable, which query the cache is keyed on. Each of those has a defensible answer here, and each has an ADR recording the alternatives that were rejected.
The request plane#
Every frontend call goes to a relative /api/… path, which the Next server rewrites onto FastAPI. Two consequences follow, and the second one is the real reason.
First, there is no CORS anywhere in the system — same-origin requests skip the mechanism entirely. That matters more than usual here because the chat response is a stream, and CORS combined with streaming is genuinely unpleasant to debug: a missing or wrong header produces a stream that simply stops, with no useful error in either the browser console or the server log.
Second, one environment variable covers two environments. In Docker the backend is api:8000 on the compose network; locally it is localhost:8000. Only API_ORIGIN changes, it is read on the Next server, and it never reaches the browser — which has no idea the backend exists as a separate thing.
Storage layout#
Five tables and one cache, with the versioning columns on documents and the tenant column denormalised onto chunks so that the isolation predicate never needs a join.
tenants
id text primary key
documents
id uuid primary key
tenant_id text references tenants
source_type text -- docs | changelog | ticket
source_path text
doc_version text
effective_date date
content_hash text -- dedup: re-ingesting identical content is a no-op
is_current boolean -- superseded documents are archived, never deleted
chunks
id uuid primary key
document_id uuid references documents
tenant_id text -- DENORMALISED: the isolation predicate never joins
chunk_index int
content text -- heading path is prepended into this (ADR-004)
heading_path text
metadata jsonb -- doc_version, product_area, error_code, conflicts_with_entry…
embedding vector(384) -- HNSW index
tsv tsvector -- GENERATED from content; GIN index
is_current boolean
unique (document_id, chunk_index)
embedding_cache -- sha256(model + text) -> vector. Makes re-ingestion nearly free.
traces -- one row per request, written after the response
escalations -- one row per abstention, with history and the top 10 chunks + scores
feedback -- thumbs, keyed by trace_idWhy tenant_id is denormalised onto chunks
Because the isolation predicate is on the hot path of every single read, and a join is a thing a developer can forget to write. With tenant_id on the row, the scope that owns the FROM clause can weld the predicate on unconditionally, and the runtime tripwire that re-checks every returned row has something local to check against.
What a trace row carries, and why
One row per request, written once after the response: the action taken, the confidence and which scale it is on, the retrieved chunk ids, the citation report as JSONB with its grounding rate precomputed, provider and model, tokens, virtual cost, and per-stage latencies. Writing traces never raises — an observability failure must not become a user-facing one.
The score kind is on the row rather than inferred later because top_score = 0.02 is uninterpretable six weeks after the fact. It is a healthy fusion score and a catastrophic reranker score, and a trace that cannot be read is not observability.
Every tuning knob, and where its value came from#
The status column is the point of this table. A number that looks tuned and never was is worse than an obvious placeholder — one of them was wrong by 4x in the direction that made its own feature invisible, and it sat in configuration with a confident comment explaining its reasoning.
| Setting | Value | Provenance | Reasoning |
|---|---|---|---|
retrieval_candidates_per_leg | 20 | design | Asking each leg for more than we keep is nearly free and gives fusion room to promote a chunk one leg ranked 18th and the other 2nd |
retrieval_fusion_top_k | 20 | design | The candidate set that recall@20 is computed over |
rerank_input_top_k | 8 | measured | The cross-encoder costs ~270 ms per pair on CPU. Eight instead of twenty is ~2.5x faster and retrieval is unaffected, because fusion still emits twenty |
rerank_top_k | 5 | design | More context costs tokens, latency, and needle-in-haystack dilution |
rrf_k | 60 | literature | Cormack et al. 2009. Damps rank 1 against rank 2 to under 2%, so one leg cannot steamroll the other |
rrf_weight_bm25 / _vector | 1.0 / 1.0 | open question | Weights were avoided on principle, then evaluation measured BM25 actively hurting multi-turn recall. 0.5 is the obvious next experiment and has not been run |
hnsw_ef_search | 100 | design | pgvector's default of 40 is too narrow once a selective tenant filter discards most of the beam |
conditional_rerank_enabled | false | deliberate | Always-rerank is the quality ceiling the evaluation measures the gate against. Shipping the gate on would make the gated arm the baseline and delete the comparison |
rerank_margin_threshold | 0.10 | measured, still not firing | The original 0.30 was wrong by 4x. Real margins are 0.055–0.076, and even 0.10 never fires — so conditional reranking is implemented and not yet demonstrated to do anything on this corpus |
confidence_threshold_rerank | 0.45 | guess | Provisional and labelled so. `make tune` sweeps it against the golden set |
confidence_threshold_fused | 0.015 | guess | Separate from the above because the two scales differ by roughly 30x |
semantic_cache_threshold | 0.95 | design + guardrails | Survivable only because identifier-bearing queries skip the semantic cache entirely and abstentions are never cached |
citation_similarity_threshold | 0.50 | deliberate | Lenient on purpose — it is catching UNRELATED citations, not grading paraphrase quality |
cache_ttl_seconds | 3600 | design | Shorter than the freshness requirement. It is the backstop for bugs in active invalidation, not the primary mechanism |
The configuration reference lists every environment variable alongside these →
Build order#
The system was built in seven phases, each of which had to be demonstrably working before the next began. The ordering is not arbitrary: infrastructure and the LLM client came first so that everything after could assume a working provider chain; the evaluation harness came after the pipeline so it had something real to score, and immediately contradicted two of the pipeline's design assumptions.
| Phase | What it added | Decisions |
|---|---|---|
| 0 · Foundations | Schema and migrations, the four-provider LLM client with retries and failover, budget tracking, /health | ADR-001, 002, 003, 005, 006 |
| 1 · Ingestion | Three loaders, three chunking strategies, cached embeddings, versioning and conflict tagging | ADR-004, 008, 009, 010 |
| 2 · Retrieval | The keyword and vector legs, rank fusion, the cross-encoder, and the isolation core | ADR-011, 012, 013, 014 |
| 3 · Generation | Query rewriting, the confidence gate, grounded generation, citation validation, escalations, traces | ADR-015, 016, 017, 018 |
| 4 · Evaluation | The golden set, stable locators, the metrics, the LLM judge, the CI regression gate | ADR-019, 020, 021, 022 |
| 5 · Cache & feedback | Exact and semantic caching, active invalidation, feedback triage, the admin statistics endpoint | ADR-023, 024, 025, 026 |
| 6 · Frontend | The assistant, the operations dashboard, and this documentation site | ADR-027, 028, 029 |