Operations
Deployment
Read this first#
| Component | Resident memory |
|---|---|
| PyTorch runtime | 700 MB – 1 GB |
| bge-small-en-v1.5 (embeddings) | ~130 MB |
| bge-reranker-base (cross-encoder) | ~280 MB |
| FastAPI + asyncpg + Redis client | ~100 MB |
| Realistic total | 1.5 – 2 GB |
Path A — drop the reranker#
RERANKER_ENABLED=falseResident memory falls to roughly 1–1.2 GB, and p95 latency drops from about five seconds to well under one.
Path B — a host with real memory#
Hugging Face Spaces' free CPU tier offers 2 vCPU and 16 GB of RAM, because it is built for exactly this kind of model-loading workload. Everything runs, reranker included.
- Free Spaces sleep after 48 hours of inactivity. Waking one reloads both models — a 30–60 second first request.
- There is active community concern about the free CPU tier and Docker SDK access for unpaid accounts. Verify the current terms before depending on it.
The free-tier stack#
| Piece | Service | Why |
|---|---|---|
| Postgres + pgvector | Neon or Supabase | The whole corpus is ~30 MB — nowhere near any free-tier limit |
| Redis | Upstash | The answer cache. The app degrades gracefully without it, so exhausting the command quota is survivable |
| Backend | HF Spaces (Path B), or Fly.io / Koyeb (Path A) | Memory is the deciding factor, and it decides this row |
| Frontend | Vercel | Next.js's native host, and the rewrite proxy works unchanged |
Step by step#
1 · Database
export DATABASE_URL="postgresql://user:pass@ep-xxx.neon.tech/fishack?sslmode=require"
psql "$DATABASE_URL" -c "CREATE EXTENSION IF NOT EXISTS vector;"
python scripts/migrate.py2 · Seed from your machine, not the server
Ingestion needs the embedding model and about two minutes of CPU. Doing that on a memory-constrained host is the slowest and most fragile part of any deployment, so run it locally against the remote database — or dump and restore, which avoids re-embedding entirely.
# from your laptop, where the model already is
DATABASE_URL="postgresql://…neon.tech/fishack?sslmode=require" python scripts/ingest.py run
# or, faster
pg_dump "postgresql://fishack:fishack@localhost:5432/fishack" \
--no-owner --no-privileges -Fc -f fishack.dump
pg_restore -d "$DATABASE_URL" --no-owner --no-privileges fishack.dumpEmbeddings are deterministic and stored in the embedding cache, so the restored database is byte-identical to the local one.
3 · Redis
REDIS_URL=rediss://default:xxx@xxx.upstash.io:6379Note rediss:// — Upstash requires TLS. The free command allowance is finite and the cache spends it on every request, but every Redis call is wrapped, so exhausting the quota degrades to a permanent cache miss rather than an outage.
4 · Backend
On Spaces, use the Docker SDK and set DATABASE_URL, REDIS_URL and GROQ_API_KEY as secrets, not public variables. The Space port is 7860, so either expose that in the Dockerfile or declare the app port in the Space front-matter. On Fly.io or Koyeb, build the same Dockerfile with RERANKER_ENABLED=false and at least 1 GB.
curl https://your-backend-url/health5 · Frontend
Import the repository, set the root directory to frontend, and add one environment variable:
API_ORIGIN=https://your-backend-urlThat is all. The Next config rewrites /api/* onto that origin, so the browser still sees a single origin and there is no CORS to configure — the same reason the proxy exists locally.
What a free-tier deployment will actually feel like#
Say this out loud rather than letting someone discover it.
| Reality | |
|---|---|
| Cold start | 30–60s after the host sleeps — both models reload. An idle Neon project also suspends |
| First query | Slow even when warm, if the cache is empty |
| LLM quota | Groq's free tier has a low tokens-per-minute ceiling. Under any real concurrency the fallback chain will fire — which is the resilience pattern working, and it is visible on the dashboard |
| Cost | $0. Every figure on the dashboard is virtual cost — what the usage would cost at paid-API prices |
Keeping it awake
A scheduled ping every thirty minutes stops the Space sleeping and the database suspending. Check the host's terms first — some free tiers consider it abuse. And it does not make a deployment production-grade; it makes a demonstration reliable enough to show someone.
name: keepalive
on:
schedule: [{ cron: "*/30 * * * *" }]
jobs:
ping:
runs-on: ubuntu-latest
steps:
- run: curl -sf ${{ secrets.BACKEND_URL }}/health || trueBefore you make it public#
Deployment checklist#
- CREATE EXTENSION vector on the remote database
- python scripts/migrate.py against DATABASE_URL
- Corpus seeded — verify SELECT count(*) FROM chunks WHERE is_current returns 312
- REDIS_URL uses rediss:// (Upstash requires TLS)
- RERANKER_ENABLED=false if the host has under ~2 GB
- API keys set as secrets, never committed
- /admin protected with a token, or unreachable
- API_ORIGIN set on the frontend host
- /health returns pgvector_installed: true
- One real question answered end to end, with citations