Every constant, threshold and model name lives in one file, and each one carries a comment explaining how its value was chosen. That rule exists because a number that looks tuned and never was is worse than an obvious placeholder.
No constant appears inline in code. Environment variables map to field names case-insensitively, so DATABASE_URL becomes database_url, and .env.example documents every one.
Model names are configuration rather than code for a specific reason: free-tier lineups change monthly, and providers retire models with little notice. A 404 from a provider is a one-line fix in .env rather than a code change and a deploy.
First is primary, the rest are failover targets. Groq leads on the fastest inference and the most generous free request quota; Ollama is last because it is optional and slowest
GROQ_API_KEY / GROQ_MODEL
llama-3.1-8b-instant
The only required key
GOOGLE_API_KEY / GEMINI_MODEL
gemini-2.5-flash
Optional fallback
OPENROUTER_API_KEY / OPENROUTER_MODEL
meta-llama/llama-3.3-70b-instruct:free
Optional fallback
OLLAMA_ENABLED / OLLAMA_BASE_URL / OLLAMA_MODEL
false · localhost:11434 · llama3.1:8b
Fully offline, needs roughly 8 GB of RAM
Retries and failover
Setting
Value
Why
retry_max_attempts
3
Per provider, before failing over
retry_base_delay
1.0 s
Exponential backoff with jitter
retry_max_delay
20.0 s
Doubles as the quota-exhaustion signal: a Retry-After LARGER than this means the daily quota is spent, so the client fails over immediately instead of sleeping. The line is drawn here because this value is already defined as the longest we are ever willing to wait on one provider
llm_timeout_seconds
60.0
Per request
Free tiers signal two very different situations with the same 429. Retry-After: 2 is a burst limit and retrying the same provider is right; Retry-After: 3600 means that provider is dead for the window and no amount of waiting fixes it.
Real spend is zero, so cost is modelled: for each model in use, the paid-tier price of the same model, or the closest comparable hosted model. A conservative mid-range fallback covers anything not in the table, so cost is never silently zero for an unknown model.
app/config.py
VIRTUAL_PRICES = { # USD per 1M tokens · snapshot: July 2026
"llama-3.1-8b-instant": (0.05, 0.08),
"llama-3.3-70b-versatile": (0.59, 0.79),
"qwen/qwen3-32b": (0.29, 0.59),
"gemini-2.5-flash": (0.30, 2.50),
"…:free": (0.59, 0.79), # priced as the paid route
"llama3.1:8b": (0.05, 0.08), # local — priced as a comparable hosted 8B
}
DEFAULT_VIRTUAL_PRICE = (0.50, 1.50)
With paid APIs this table disappears — cost comes from the provider invoice and the tracking code stays identical.