Skip to content

Operations

Configuration

Every constant, threshold and model name lives in one file, and each one carries a comment explaining how its value was chosen. That rule exists because a number that looks tuned and never was is worse than an obvious placeholder.

The rule#

No constant appears inline in code. Environment variables map to field names case-insensitively, so DATABASE_URL becomes database_url, and .env.example documents every one.

Model names are configuration rather than code for a specific reason: free-tier lineups change monthly, and providers retire models with little notice. A 404 from a provider is a one-line fix in .env rather than a code change and a deploy.

Infrastructure#

VariableDefaultNotes
DATABASE_URLpostgresql://fishack:fishack@localhost:5432/fishackNeeds the pgvector extension enabled
REDIS_URLredis://localhost:6379/0Use rediss:// for hosted Redis — Upstash requires TLS
ADMIN_TOKEN(empty)Empty means the statistics endpoint is OPEN. Set it before anything is public
API_ORIGINhttp://localhost:8000Frontend only. Read on the Next server, never shipped to the browser

LLM providers#

VariableDefaultNotes
LLM_PROVIDER_ORDERgroq,gemini,openrouter,ollamaFirst is primary, the rest are failover targets. Groq leads on the fastest inference and the most generous free request quota; Ollama is last because it is optional and slowest
GROQ_API_KEY / GROQ_MODELllama-3.1-8b-instantThe only required key
GOOGLE_API_KEY / GEMINI_MODELgemini-2.5-flashOptional fallback
OPENROUTER_API_KEY / OPENROUTER_MODELmeta-llama/llama-3.3-70b-instruct:freeOptional fallback
OLLAMA_ENABLED / OLLAMA_BASE_URL / OLLAMA_MODELfalse · localhost:11434 · llama3.1:8bFully offline, needs roughly 8 GB of RAM

Retries and failover

SettingValueWhy
retry_max_attempts3Per provider, before failing over
retry_base_delay1.0 sExponential backoff with jitter
retry_max_delay20.0 sDoubles as the quota-exhaustion signal: a Retry-After LARGER than this means the daily quota is spent, so the client fails over immediately instead of sleeping. The line is drawn here because this value is already defined as the longest we are ever willing to wait on one provider
llm_timeout_seconds60.0Per request
Free tiers signal two very different situations with the same 429. Retry-After: 2 is a burst limit and retrying the same provider is right; Retry-After: 3600 means that provider is dead for the window and no amount of waiting fixes it.

Retrieval and ranking#

SettingValueProvenance
retrieval_candidates_per_leg20design
retrieval_fusion_top_k20design
rerank_input_top_k8measured
rerank_top_k5design
rrf_k60literature
rrf_weight_bm25 / _vector1.0 / 1.0open question
hnsw_ef_search100design
embedding_model_nameBAAI/bge-small-en-v1.5design
embedding_dimchanging it needs a column migration, a reindex AND a full re-ingest384hard constraint
RERANKER_ENABLEDtruedeployment lever
reranker_model_nameBAAI/bge-reranker-basedesign
reranker_batch_size16design
reranker_max_length512model limit
conditional_rerank_enabledfalsedeliberate
rerank_ambiguity_window5design
rerank_margin_threshold0.10measured

Generation and validation#

SettingValueProvenance
confidence_threshold_rerank0.45guess
confidence_threshold_fused0.015guess
llm_temperature0.1design
llm_max_tokens1024design
query_rewrite_enabledtruedesign
query_rewrite_history_turns6guess
query_rewrite_max_tokens120design
citation_validation_enabledtruedesign
citation_similarity_threshold0.50deliberate
abstention_messagethe prompt, the detector and the eval assertions must all agree on itfixed stringcontract

Caching#

SettingValueWhy
CACHE_ENABLEDtrueTurn off entirely on a Redis-free deployment. Every call is wrapped anyway, so a failure degrades to a miss
cache_ttl_seconds3600Shorter than the freshness requirement. It is a backstop for bugs in active invalidation, not the primary mechanism
semantic_cache_enabledtrueThe fuzzy path
semantic_cache_threshold0.95Survivable only because identifier queries skip this path entirely and abstentions are never cached
semantic_cache_max_candidates200Bounds the similarity scan

What the provenance labels mean#

LabelMeans
designComes from the system design, and has a stated rationale
literatureFrom a published result, cited in the code comment
measuredAn experiment produced this number, and the experiment is reproducible
guessProvisional, and labelled so in the config file itself. These are the ones to distrust first
deliberateChosen against the obvious value, for a reason recorded in an ADR
open questionEvidence exists that this is wrong, and the experiment has not been run

The virtual price table#

Real spend is zero, so cost is modelled: for each model in use, the paid-tier price of the same model, or the closest comparable hosted model. A conservative mid-range fallback covers anything not in the table, so cost is never silently zero for an unknown model.

app/config.py
VIRTUAL_PRICES = {                        # USD per 1M tokens · snapshot: July 2026
    "llama-3.1-8b-instant":      (0.05, 0.08),
    "llama-3.3-70b-versatile":   (0.59, 0.79),
    "qwen/qwen3-32b":            (0.29, 0.59),
    "gemini-2.5-flash":          (0.30, 2.50),
    "…:free":                    (0.59, 0.79),   # priced as the paid route
    "llama3.1:8b":               (0.05, 0.08),   # local — priced as a comparable hosted 8B
}
DEFAULT_VIRTUAL_PRICE = (0.50, 1.50)

With paid APIs this table disappears — cost comes from the provider invoice and the tracking code stays identical.

The decision record →