Operations
API reference
Chat#
/chatAnswer a question, streaming.Request
{
"tenant_id": "acme",
"query": "What is the webhook retry limit?",
"messages": [
{ "role": "user", "content": "…" },
{ "role": "assistant", "content": "…" }
],
"conversation_id": "b6f1…" // client-generated UUID, or null
}| Field | Type | Notes |
|---|---|---|
tenant_id | string | Validated against the tenants table, so a typo returns 404 rather than an empty result set that reads as 'we have no documentation about that' |
query | string | Empty or whitespace-only returns 422 |
messages | array | Prior turns, text only. The server is stateless — history is an INPUT to query rewriting, not server state, which is why any replica can serve any request |
conversation_id | string | null | Persisted on the trace row so multi-turn traces group together in observability, feedback triage and the golden set. Not used for state |
Response — server-sent events, in this order
event: meta
data: {"data":{"citations":[…],"gate":{…},"rewrite":{…},"cache_status":"miss"}}
event: delta
data: {"text":"The webhook retry limit is "}
event: delta
data: {"text":"5 attempts [2]. "}
event: final
data: {"data":{ … the complete ChatResponse … }}meta arriving before the first delta is a contract, not an implementation detail: it is what lets the sources panel populate while the answer is still being written. An error event only ever appears after the stream has begun — before that, errors are ordinary HTTP, because once the status line has been sent a 500 is no longer available.
The final payload
| Field | Type | Notes |
|---|---|---|
answer | string | The full text, with inline [n] markers |
action | "answered" | "abstained" | "escalated" | "cache_hit" | All three abstention paths set 'escalated', so the escalation rate cannot understate reality |
citations[] | Citation[] | index, chunk_id, title, source_type, source_path, heading_path, doc_version, effective_date, score, is_contested, was_cited |
citation_report | CitationReport | null | claims[] with per-claim supported/similarity, invalid_indices (fabrications), unused_indices |
gate | GateDecision | null | should_generate, reason, top_score, threshold, and score_kind — the last is required to interpret the first two, because the scales differ ~30x |
rewrite | RewriteResult | null | original, rewritten, changed, skipped_reason, elapsed_ms |
cache_status | "miss" | "exact_hit" | "semantic_hit" | With cache_similarity on a semantic hit — the reader is looking at an answer written for a different question |
virtual_cost_usd | number | Zero on a cache hit, deliberately. Replaying the original spend would make cost rise as caching improved |
*_ms | number | rewrite, retrieval, rerank, generation, validation, total. Zero means the stage did not run |
degraded_legs[] | string[] | A retrieval leg that failed. The answer was built on less evidence than usual |
/chat/syncThe same pipeline, one JSON object, no streaming.It exists because streaming is miserable to test with a plain HTTP client and because the evaluation harness wants a single object. It runs the identical pipeline — the synchronous entry point simply drains the stream — so it can never diverge from what users actually get.
Feedback#
/feedbackRecord a thumbs rating against a trace.{ "trace_id": "…", "rating": 1, "comment": null } // rating: 1 | -1Fire-and-forget from the client's point of view: the button lights optimistically, and a failed POST is not worth interrupting someone's reading over. The feedback flywheel loses one data point and nothing else breaks.
Admin#
/admin/stats?hours=&tenant_id=Aggregates from the traces table. Reads across tenants.Guarded by an admin-token dependency on the router. Every figure is computed by the backend from the traces table; the frontend does no arithmetic beyond formatting, deliberately — if the dashboard derived its own rates they could disagree with what the eval harness and the triage script read from the same rows.
{
"window": { "hours": 24, "since": "…", "tenant_id": null },
"requests": { "total": 0, "escalation_rate": 0, "cache_hit_rate": 0, "by_action": {} },
"latency_ms": { "p50": 0, "p95": 0, "mean_retrieval": 0, "mean_rerank": 0, "mean_generation": 0 },
"cost": { "methodology": "…", "total_usd": 0, "per_query_usd": 0, "tokens_in": 0, "tokens_out": 0 },
"quality": { "thumbs_up": 0, "thumbs_down": 0, "satisfaction_rate": 0,
"answers_with_fabricated_citations": 0, "mean_confidence": 0 },
"escalations": { "open": 0, "by_reason": {} },
"by_tenant": [ { "tenant_id": "acme", "requests": 0, … } ],
"providers": { "groq": { "requests": 0, "tokens_in": 0, "virtual_cost_usd": 0 } }
}/admin/cache/flush?tenant_id=Drop one tenant's cached answers.Health#
/healthLiveness plus an independent report per dependency.Each dependency reports separately so a broken Redis does not mask a healthy Postgres. The endpoint reports rather than raises — a health check that can fail is one more thing to page about.
Error taxonomy#
| Status | When | Shape |
|---|---|---|
| 422 | Empty query | FastAPI validation detail |
| 404 | Unknown tenant | detail: unknown tenant 'x' — deliberately not an empty result set |
| 401 | Admin token missing or wrong | The dashboard says which side to fix |
| 502 | The frontend route handler cannot reach the API | detail naming the origin it tried |
| 200 + error event | The pipeline raises after the stream has begun | A terminal SSE error event. The client gets a definite end state rather than a connection that just stops, which is indistinguishable from a network failure |