Operations
Running and operating
The four questions#
Every panel on the operations dashboard answers one of these. Naming the sections after the decision rather than the metric is the whole reorganisation: “Latency” is a metric family, “Is it fast?” is what someone opened the page to find out.
| Question | What it shows | Which direction is bad |
|---|---|---|
| Is it working? | Request volume, escalation rate, cache hit rate, cost per query, and the outcome of every request | Escalation rate cuts BOTH ways — see below |
| Is it fast? | p50, p95, and mean latency per stage as bars | p95 above 3s. The bars say which stage owns it, and it is almost always the reranker |
| Is it trustworthy? | Satisfaction, fabricated-citation count, mean confidence, open escalations by reason | Any fabricated citation at all. Satisfaction is noisy at low volume |
| Is it leaking? | Per-tenant traffic, escalation rate, cache hits, p95 and cost | A row appearing that should not exist. The numbers themselves are just volume |
| What would it cost? | Virtual cost totals, token counts, and per-provider usage against free-tier quota | One provider carrying everything means the fallback chain is untested in practice |
Escalation rate is the one metric that cuts both ways
Too high means retrieval is failing, or the corpus has a real gap, or the threshold is miscalibrated. Too low means the confidence gate is protecting nobody — a system that never declines is a system that answers questions it cannot answer, which is the exact failure the whole design exists to prevent.
So it is displayed with a good/warn/bad band rather than as “lower is better”, and the breakdown by reason is what makes it actionable.
| Escalation reason | What a spike means | What to do |
|---|---|---|
low_confidence | Retrieval found something, and not strongly enough | Run `make tune` against the golden set; check whether recent ingestion changed chunk boundaries |
no_results | Nothing matched at all — usually genuinely out of scope | Read the escalation queries. A cluster is a documentation gap, and it is the clearest signal you will ever get about what customers ask that you have not written down |
model_abstained | Context was topically right and factually silent | Also a corpus gap, and a more interesting one — retrieval found the right area and the answer was not there |
generation_failed | Every provider in the chain errored or was exhausted | Check provider quota on the same page. A spike here beside a flat low_confidence count is an outage signal, not a quality signal |
Reading the numbers honestly#
| Metric | What would make it misleading |
|---|---|
| Satisfaction | Computed over RATED answers only. At demo volume a single vote moves it several points — read the direction, not the decimal |
| Mean confidence | Averages two scales roughly 30x apart, because some requests were reranked and some were not. It is a trend, never an absolute — the per-answer pill in the assistant shows which scale each one is on |
| Per-stage latency | Means, not percentiles — per-stage percentiles are not recorded, so one slow request moves a bar in a way it would not move a p95. The shape is the finding |
| Cache hit rate | Includes semantic hits, which serve an answer written for a DIFFERENT question. A high rate is good; a high SEMANTIC rate deserves a look at what is matching |
| Cost per query | Virtual. And a cache hit records zero — deliberately, because replaying the original spend would make cost rise as caching improved |
Health and startup#
GET /health reports each dependency independently, so a broken Redis does not mask a healthy Postgres or the other way round. It never raises — a health endpoint that can fail is one more thing to page about.
{
"status": "ok",
"database": { "ok": true, "pgvector_installed": true },
"redis": { "ok": true },
"llm_chain": ["groq", "gemini", "openrouter"]
}Startup loads both local models before the app accepts traffic, and fails the boot if the embedding model's dimension does not match the schema's vector(384). A dimension mismatch that survives startup would surface as a per-query error under load, which is a far worse place to find out.
Feedback triage#
Thumbs-down feedback is classified by heuristic rather than by a model, because every signal needed is already on the trace row: the confidence and its scale, the citation report, whether a cited chunk was contested, and the cache status. So classification is free, instant, deterministic and explainable — “retrieval failure, because confidence was 0.31 against a 0.45 threshold” is a sentence you can argue with.
The check order is the actual design
| Order | Check | Why it comes here |
|---|---|---|
| 1 | Cache | A hit means retrieval and generation never ran for this request. Blaming retrieval for an answer it did not produce sends someone to debug the wrong component entirely |
| 2 | Stale data | The model may have followed the prompt perfectly while the context was out of date. That is an ingestion problem wearing a generation problem's clothes |
| 3 | Retrieval | If the right chunks never arrived, no prompt could have produced a good answer |
| 4 | Generation | What is left once the three above are ruled out |
unclear is a real category and is reported separately rather than forced into a neighbour. A misclassified failure is worse than an unclassified one — it points at the wrong component with confidence, and the real bug survives the investigation.
Runbook#
| Symptom | Most likely cause | First action |
|---|---|---|
| Escalation rate jumps, low_confidence dominates | A re-ingest changed chunk boundaries, or the threshold drifted from the corpus | Run `make eval-retrieval` and compare against the committed baseline |
| generation_failed spikes | Every provider exhausted, or a model name was retired | Check provider usage on the dashboard, then /health. A 404 from a provider is a config problem |
| Cache hit rate falls to zero | Redis is unreachable, or its free-tier command quota is spent | Nothing breaks — every call degrades to a miss. Confirm on /health, then decide whether to disable the cache explicitly |
| p95 above 3s | The cross-encoder | Set RERANKER_ENABLED=false. Measurements say vector-only has the better recall@5 on this corpus anyway |
| A fabricated citation appears | The model invented a marker index | Read the trace. This is already surfaced in the answer and counted on the dashboard — it should be zero, and any non-zero value is worth a look |
| degraded_legs on many answers | One retrieval leg is erroring | Answers are still being built, on less evidence. Check Postgres and the GIN/HNSW indexes |
| A tenant appears in the per-tenant table that should not exist | This is the leak signal | Stop and investigate. The isolation tripwire should have raised before this — if it did not, a code path is bypassing the scope |
Protecting the dashboard#
Pick one before deploying:
- Set an admin token. The frontend's route handler attaches it server-side, so it never reaches the browser bundle. Same value on both sides.
- Block it at the edge. Remove the rewrite for
/api/admin/*and it is unreachable through the frontend. - Do not deploy the dashboard and read statistics through a tunnel.
In production this belongs behind SSO with an admin role, with every access audited. A shared token is the floor, not the target.