Skip to content

Operations

Running and operating

The dashboard is organised around the four questions an operator actually arrives with, not around the metric families the API happens to return. This page says what each number means, which direction is bad, and what to do when it moves.

The four questions#

Every panel on the operations dashboard answers one of these. Naming the sections after the decision rather than the metric is the whole reorganisation: “Latency” is a metric family, “Is it fast?” is what someone opened the page to find out.

QuestionWhat it showsWhich direction is bad
Is it working?Request volume, escalation rate, cache hit rate, cost per query, and the outcome of every requestEscalation rate cuts BOTH ways — see below
Is it fast?p50, p95, and mean latency per stage as barsp95 above 3s. The bars say which stage owns it, and it is almost always the reranker
Is it trustworthy?Satisfaction, fabricated-citation count, mean confidence, open escalations by reasonAny fabricated citation at all. Satisfaction is noisy at low volume
Is it leaking?Per-tenant traffic, escalation rate, cache hits, p95 and costA row appearing that should not exist. The numbers themselves are just volume
What would it cost?Virtual cost totals, token counts, and per-provider usage against free-tier quotaOne provider carrying everything means the fallback chain is untested in practice

Escalation rate is the one metric that cuts both ways

Too high means retrieval is failing, or the corpus has a real gap, or the threshold is miscalibrated. Too low means the confidence gate is protecting nobody — a system that never declines is a system that answers questions it cannot answer, which is the exact failure the whole design exists to prevent.

So it is displayed with a good/warn/bad band rather than as “lower is better”, and the breakdown by reason is what makes it actionable.

Escalation reasonWhat a spike meansWhat to do
low_confidenceRetrieval found something, and not strongly enoughRun `make tune` against the golden set; check whether recent ingestion changed chunk boundaries
no_resultsNothing matched at all — usually genuinely out of scopeRead the escalation queries. A cluster is a documentation gap, and it is the clearest signal you will ever get about what customers ask that you have not written down
model_abstainedContext was topically right and factually silentAlso a corpus gap, and a more interesting one — retrieval found the right area and the answer was not there
generation_failedEvery provider in the chain errored or was exhaustedCheck provider quota on the same page. A spike here beside a flat low_confidence count is an outage signal, not a quality signal

Reading the numbers honestly#

MetricWhat would make it misleading
SatisfactionComputed over RATED answers only. At demo volume a single vote moves it several points — read the direction, not the decimal
Mean confidenceAverages two scales roughly 30x apart, because some requests were reranked and some were not. It is a trend, never an absolute — the per-answer pill in the assistant shows which scale each one is on
Per-stage latencyMeans, not percentiles — per-stage percentiles are not recorded, so one slow request moves a bar in a way it would not move a p95. The shape is the finding
Cache hit rateIncludes semantic hits, which serve an answer written for a DIFFERENT question. A high rate is good; a high SEMANTIC rate deserves a look at what is matching
Cost per queryVirtual. And a cache hit records zero — deliberately, because replaying the original spend would make cost rise as caching improved

Health and startup#

GET /health reports each dependency independently, so a broken Redis does not mask a healthy Postgres or the other way round. It never raises — a health endpoint that can fail is one more thing to page about.

json
{
  "status": "ok",
  "database": { "ok": true, "pgvector_installed": true },
  "redis": { "ok": true },
  "llm_chain": ["groq", "gemini", "openrouter"]
}

Startup loads both local models before the app accepts traffic, and fails the boot if the embedding model's dimension does not match the schema's vector(384). A dimension mismatch that survives startup would surface as a per-query error under load, which is a far worse place to find out.

Feedback triage#

Thumbs-down feedback is classified by heuristic rather than by a model, because every signal needed is already on the trace row: the confidence and its scale, the citation report, whether a cited chunk was contested, and the cache status. So classification is free, instant, deterministic and explainable — “retrieval failure, because confidence was 0.31 against a 0.45 threshold” is a sentence you can argue with.

The check order is the actual design

OrderCheckWhy it comes here
1CacheA hit means retrieval and generation never ran for this request. Blaming retrieval for an answer it did not produce sends someone to debug the wrong component entirely
2Stale dataThe model may have followed the prompt perfectly while the context was out of date. That is an ingestion problem wearing a generation problem's clothes
3RetrievalIf the right chunks never arrived, no prompt could have produced a good answer
4GenerationWhat is left once the three above are ruled out

unclear is a real category and is reported separately rather than forced into a neighbour. A misclassified failure is worse than an unclassified one — it points at the wrong component with confidence, and the real bug survives the investigation.

Runbook#

SymptomMost likely causeFirst action
Escalation rate jumps, low_confidence dominatesA re-ingest changed chunk boundaries, or the threshold drifted from the corpusRun `make eval-retrieval` and compare against the committed baseline
generation_failed spikesEvery provider exhausted, or a model name was retiredCheck provider usage on the dashboard, then /health. A 404 from a provider is a config problem
Cache hit rate falls to zeroRedis is unreachable, or its free-tier command quota is spentNothing breaks — every call degrades to a miss. Confirm on /health, then decide whether to disable the cache explicitly
p95 above 3sThe cross-encoderSet RERANKER_ENABLED=false. Measurements say vector-only has the better recall@5 on this corpus anyway
A fabricated citation appearsThe model invented a marker indexRead the trace. This is already surfaced in the answer and counted on the dashboard — it should be zero, and any non-zero value is worth a look
degraded_legs on many answersOne retrieval leg is erroringAnswers are still being built, on less evidence. Check Postgres and the GIN/HNSW indexes
A tenant appears in the per-tenant table that should not existThis is the leak signalStop and investigate. The isolation tripwire should have raised before this — if it did not, a code path is bypassing the scope

Protecting the dashboard#

Pick one before deploying:

  • Set an admin token. The frontend's route handler attaches it server-side, so it never reaches the browser bundle. Same value on both sides.
  • Block it at the edge. Remove the rewrite for /api/admin/* and it is unreachable through the frontend.
  • Do not deploy the dashboard and read statistics through a tunnel.

In production this belongs behind SSO with an admin role, with every access audited. A shared token is the floor, not the target.

Deployment, where memory is the whole problem →