Start here
Quickstart
Three commands bring up Postgres, Redis, the API and the UI together, and ingest a committed corpus. The whole path needs exactly one free API key, and the first run is deliberately slow.
Three commands#
shell
cp .env.example .env # add a free Groq key — the rest are optional fallbacks
docker compose up -d --build
docker compose exec api python scripts/ingest.py runThen open http://localhost:3000 for the documentation and the assistant, and http://localhost:3000/admin for operations.
What each step is actually doing#
One API key, and which one
.envOnly one provider is required. Groq is the recommended primary — fastest inference and the most generous free daily request quota of the four. Gemini, OpenRouter and a local Ollama are optional fallbacks; the chain fails over automatically when one is rate-limited or exhausted.First boot downloads two models, on purpose
~450 MBThe API container fetchesbge-small-en-v1.5(~130 MB) andbge-reranker-base(~280 MB) at startup, not on first request. Paying that during boot means the first user gets a normal response instead of a ten-second one, and a missing model fails the deploy rather than the first customer. Watch forFishack up. LLM chain: groq -> …indocker compose logs -f api.Ingestion embeds the corpus
~312 chunks · ~2 min156 chunks per tenant across two tenants, embedded on CPU. The step is idempotent — documents are deduplicated by content hash, so re-running is nearly free. Verify withscripts/ingest.py stats.Confirm before building anything on top
369 testsscripts/smoke_test.pychecks Postgres, the pgvector extension, Redis and every configured provider with a real call, so a misconfigured provider surfaces here rather than three layers later.pytest -q -m "not integration"should report 369 passed.
Providers and their free tiers#
| Provider | Where to get a key | Free tier |
|---|---|---|
| Groqprimary — first in the chain | console.groq.com → API Keys | Generous daily request quota, low tokens per minute |
| Gemini | aistudio.google.com | Tight daily quota; good as a second hop |
| OpenRouter | openrouter.ai | `:free` model routes |
| Ollama | local install | Fully offline, needs roughly 8 GB of RAM |
Local development without Docker#
A faster loop, with hot reload on both sides. Infrastructure stays in Docker; the API and the frontend run on the host.
shell
docker compose up -d postgres redis # infrastructure only
python -m venv .venv
.venv/Scripts/activate # or: source .venv/bin/activate
pip install -e ".[dev]"
python scripts/migrate.py
python scripts/ingest.py run
make api # terminal 1
cd frontend && npm install && npm run dev # terminal 2The commands worth knowing#
| Command | What it does |
|---|---|
make test | 369 unit tests plus 23 integration tests |
make eval-retrieval | Retrieval scorecard across five arms. No LLM calls, under a minute, fully reproducible. |
make eval | The full harness including LLM-as-judge, compared against the committed baseline |
make chunking-experiment | Naive fixed-window versus per-source chunking, on shadow tenants |
make tune | Sweep the confidence gate against the golden set |
make playground | BM25, vector, hybrid and reranked results side by side for one query |
make chat | The pipeline in a terminal, with every stage's decision printed |
make show-prompt | The exact messages the model receives |
make triage | Classify thumbs-down feedback into retrieval, generation or stale-data failures |
Next
The guided tour walks five queries in a fixed order, each one turning on a defence the previous one did not need. It is the fastest way to see what the system does, and it takes about three minutes.