Skip to content

Operations

Deployment

Memory is the whole problem, and every hosting choice below follows from it. This application runs two local transformer models, which is unusual for a web service and does not fit in the 512 MB most free tiers offer.

Read this first#

ComponentResident memory
PyTorch runtime700 MB – 1 GB
bge-small-en-v1.5 (embeddings)~130 MB
bge-reranker-base (cross-encoder)~280 MB
FastAPI + asyncpg + Redis client~100 MB
Realistic total1.5 – 2 GB

Path A — drop the reranker#

shell
RERANKER_ENABLED=false

Resident memory falls to roughly 1–1.2 GB, and p95 latency drops from about five seconds to well under one.

Path B — a host with real memory#

Hugging Face Spaces' free CPU tier offers 2 vCPU and 16 GB of RAM, because it is built for exactly this kind of model-loading workload. Everything runs, reranker included.

  • Free Spaces sleep after 48 hours of inactivity. Waking one reloads both models — a 30–60 second first request.
  • There is active community concern about the free CPU tier and Docker SDK access for unpaid accounts. Verify the current terms before depending on it.

The free-tier stack#

PieceServiceWhy
Postgres + pgvectorNeon or SupabaseThe whole corpus is ~30 MB — nowhere near any free-tier limit
RedisUpstashThe answer cache. The app degrades gracefully without it, so exhausting the command quota is survivable
BackendHF Spaces (Path B), or Fly.io / Koyeb (Path A)Memory is the deciding factor, and it decides this row
FrontendVercelNext.js's native host, and the rewrite proxy works unchanged

Step by step#

1 · Database

shell
export DATABASE_URL="postgresql://user:pass@ep-xxx.neon.tech/fishack?sslmode=require"
psql "$DATABASE_URL" -c "CREATE EXTENSION IF NOT EXISTS vector;"
python scripts/migrate.py

2 · Seed from your machine, not the server

Ingestion needs the embedding model and about two minutes of CPU. Doing that on a memory-constrained host is the slowest and most fragile part of any deployment, so run it locally against the remote database — or dump and restore, which avoids re-embedding entirely.

shell
# from your laptop, where the model already is
DATABASE_URL="postgresql://…neon.tech/fishack?sslmode=require" python scripts/ingest.py run

# or, faster
pg_dump "postgresql://fishack:fishack@localhost:5432/fishack" \
  --no-owner --no-privileges -Fc -f fishack.dump
pg_restore -d "$DATABASE_URL" --no-owner --no-privileges fishack.dump

Embeddings are deterministic and stored in the embedding cache, so the restored database is byte-identical to the local one.

3 · Redis

shell
REDIS_URL=rediss://default:xxx@xxx.upstash.io:6379

Note rediss:// — Upstash requires TLS. The free command allowance is finite and the cache spends it on every request, but every Redis call is wrapped, so exhausting the quota degrades to a permanent cache miss rather than an outage.

4 · Backend

On Spaces, use the Docker SDK and set DATABASE_URL, REDIS_URL and GROQ_API_KEY as secrets, not public variables. The Space port is 7860, so either expose that in the Dockerfile or declare the app port in the Space front-matter. On Fly.io or Koyeb, build the same Dockerfile with RERANKER_ENABLED=false and at least 1 GB.

shell
curl https://your-backend-url/health

5 · Frontend

Import the repository, set the root directory to frontend, and add one environment variable:

shell
API_ORIGIN=https://your-backend-url

That is all. The Next config rewrites /api/* onto that origin, so the browser still sees a single origin and there is no CORS to configure — the same reason the proxy exists locally.

What a free-tier deployment will actually feel like#

Say this out loud rather than letting someone discover it.

Reality
Cold start30–60s after the host sleeps — both models reload. An idle Neon project also suspends
First querySlow even when warm, if the cache is empty
LLM quotaGroq's free tier has a low tokens-per-minute ceiling. Under any real concurrency the fallback chain will fire — which is the resilience pattern working, and it is visible on the dashboard
Cost$0. Every figure on the dashboard is virtual cost — what the usage would cost at paid-API prices

Keeping it awake

A scheduled ping every thirty minutes stops the Space sleeping and the database suspending. Check the host's terms first — some free tiers consider it abuse. And it does not make a deployment production-grade; it makes a demonstration reliable enough to show someone.

.github/workflows/keepalive.yml
name: keepalive
on:
  schedule: [{ cron: "*/30 * * * *" }]
jobs:
  ping:
    runs-on: ubuntu-latest
    steps:
      - run: curl -sf ${{ secrets.BACKEND_URL }}/health || true

Before you make it public#

Deployment checklist#

  • CREATE EXTENSION vector on the remote database
  • python scripts/migrate.py against DATABASE_URL
  • Corpus seeded — verify SELECT count(*) FROM chunks WHERE is_current returns 312
  • REDIS_URL uses rediss:// (Upstash requires TLS)
  • RERANKER_ENABLED=false if the host has under ~2 GB
  • API keys set as secrets, never committed
  • /admin protected with a token, or unreachable
  • API_ORIGIN set on the frontend host
  • /health returns pgvector_installed: true
  • One real question answered end to end, with citations

The API reference →