Skip to content

Start here

Quickstart

Three commands bring up Postgres, Redis, the API and the UI together, and ingest a committed corpus. The whole path needs exactly one free API key, and the first run is deliberately slow.

Three commands#

shell
cp .env.example .env    # add a free Groq key — the rest are optional fallbacks
docker compose up -d --build
docker compose exec api python scripts/ingest.py run

Then open http://localhost:3000 for the documentation and the assistant, and http://localhost:3000/admin for operations.

What each step is actually doing#

  1. One API key, and which one

    .env
    Only one provider is required. Groq is the recommended primary — fastest inference and the most generous free daily request quota of the four. Gemini, OpenRouter and a local Ollama are optional fallbacks; the chain fails over automatically when one is rate-limited or exhausted.
  2. First boot downloads two models, on purpose

    ~450 MB
    The API container fetches bge-small-en-v1.5 (~130 MB) and bge-reranker-base (~280 MB) at startup, not on first request. Paying that during boot means the first user gets a normal response instead of a ten-second one, and a missing model fails the deploy rather than the first customer. Watch for Fishack up. LLM chain: groq -> … in docker compose logs -f api.
  3. Ingestion embeds the corpus

    ~312 chunks · ~2 min
    156 chunks per tenant across two tenants, embedded on CPU. The step is idempotent — documents are deduplicated by content hash, so re-running is nearly free. Verify with scripts/ingest.py stats.
  4. Confirm before building anything on top

    369 tests
    scripts/smoke_test.py checks Postgres, the pgvector extension, Redis and every configured provider with a real call, so a misconfigured provider surfaces here rather than three layers later. pytest -q -m "not integration" should report 369 passed.

Providers and their free tiers#

ProviderWhere to get a keyFree tier
Groqprimary — first in the chainconsole.groq.com → API KeysGenerous daily request quota, low tokens per minute
Geminiaistudio.google.comTight daily quota; good as a second hop
OpenRouteropenrouter.ai`:free` model routes
Ollamalocal installFully offline, needs roughly 8 GB of RAM
Chain order is configuration, not code: LLM_PROVIDER_ORDER defaults to groq,gemini,openrouter,ollama.

Local development without Docker#

A faster loop, with hot reload on both sides. Infrastructure stays in Docker; the API and the frontend run on the host.

shell
docker compose up -d postgres redis          # infrastructure only

python -m venv .venv
.venv/Scripts/activate                       # or: source .venv/bin/activate
pip install -e ".[dev]"
python scripts/migrate.py
python scripts/ingest.py run

make api                                     # terminal 1
cd frontend && npm install && npm run dev     # terminal 2

The commands worth knowing#

CommandWhat it does
make test369 unit tests plus 23 integration tests
make eval-retrievalRetrieval scorecard across five arms. No LLM calls, under a minute, fully reproducible.
make evalThe full harness including LLM-as-judge, compared against the committed baseline
make chunking-experimentNaive fixed-window versus per-source chunking, on shadow tenants
make tuneSweep the confidence gate against the golden set
make playgroundBM25, vector, hybrid and reranked results side by side for one query
make chatThe pipeline in a terminal, with every stage's decision printed
make show-promptThe exact messages the model receives
make triageClassify thumbs-down feedback into retrieval, generation or stale-data failures

Next

The guided tour walks five queries in a fixed order, each one turning on a defence the previous one did not need. It is the fastest way to see what the system does, and it takes about three minutes.