Skip to content

Operations

API reference

Five endpoints. The interesting one streams server-sent events in a fixed order that the interface depends on, and the second most interesting one exists only because streaming is miserable to test with a plain HTTP client.

Chat#

POST/chatAnswer a question, streaming.

Request

json
{
  "tenant_id": "acme",
  "query": "What is the webhook retry limit?",
  "messages": [
    { "role": "user", "content": "…" },
    { "role": "assistant", "content": "…" }
  ],
  "conversation_id": "b6f1…"      // client-generated UUID, or null
}
FieldTypeNotes
tenant_idstringValidated against the tenants table, so a typo returns 404 rather than an empty result set that reads as 'we have no documentation about that'
querystringEmpty or whitespace-only returns 422
messagesarrayPrior turns, text only. The server is stateless — history is an INPUT to query rewriting, not server state, which is why any replica can serve any request
conversation_idstring | nullPersisted on the trace row so multi-turn traces group together in observability, feedback triage and the golden set. Not used for state

Response — server-sent events, in this order

text
event: meta
data: {"data":{"citations":[…],"gate":{…},"rewrite":{…},"cache_status":"miss"}}

event: delta
data: {"text":"The webhook retry limit is "}

event: delta
data: {"text":"5 attempts [2]. "}

event: final
data: {"data":{ … the complete ChatResponse … }}

meta arriving before the first delta is a contract, not an implementation detail: it is what lets the sources panel populate while the answer is still being written. An error event only ever appears after the stream has begun — before that, errors are ordinary HTTP, because once the status line has been sent a 500 is no longer available.

The final payload

FieldTypeNotes
answerstringThe full text, with inline [n] markers
action"answered" | "abstained" | "escalated" | "cache_hit"All three abstention paths set 'escalated', so the escalation rate cannot understate reality
citations[]Citation[]index, chunk_id, title, source_type, source_path, heading_path, doc_version, effective_date, score, is_contested, was_cited
citation_reportCitationReport | nullclaims[] with per-claim supported/similarity, invalid_indices (fabrications), unused_indices
gateGateDecision | nullshould_generate, reason, top_score, threshold, and score_kind — the last is required to interpret the first two, because the scales differ ~30x
rewriteRewriteResult | nulloriginal, rewritten, changed, skipped_reason, elapsed_ms
cache_status"miss" | "exact_hit" | "semantic_hit"With cache_similarity on a semantic hit — the reader is looking at an answer written for a different question
virtual_cost_usdnumberZero on a cache hit, deliberately. Replaying the original spend would make cost rise as caching improved
*_msnumberrewrite, retrieval, rerank, generation, validation, total. Zero means the stage did not run
degraded_legs[]string[]A retrieval leg that failed. The answer was built on less evidence than usual
POST/chat/syncThe same pipeline, one JSON object, no streaming.

It exists because streaming is miserable to test with a plain HTTP client and because the evaluation harness wants a single object. It runs the identical pipeline — the synchronous entry point simply drains the stream — so it can never diverge from what users actually get.

Feedback#

POST/feedbackRecord a thumbs rating against a trace.
json
{ "trace_id": "…", "rating": 1, "comment": null }   // rating: 1 | -1

Fire-and-forget from the client's point of view: the button lights optimistically, and a failed POST is not worth interrupting someone's reading over. The feedback flywheel loses one data point and nothing else breaks.

Admin#

GET/admin/stats?hours=&tenant_id=Aggregates from the traces table. Reads across tenants.

Guarded by an admin-token dependency on the router. Every figure is computed by the backend from the traces table; the frontend does no arithmetic beyond formatting, deliberately — if the dashboard derived its own rates they could disagree with what the eval harness and the triage script read from the same rows.

json
{
  "window":   { "hours": 24, "since": "…", "tenant_id": null },
  "requests": { "total": 0, "escalation_rate": 0, "cache_hit_rate": 0, "by_action": {} },
  "latency_ms": { "p50": 0, "p95": 0, "mean_retrieval": 0, "mean_rerank": 0, "mean_generation": 0 },
  "cost":     { "methodology": "…", "total_usd": 0, "per_query_usd": 0, "tokens_in": 0, "tokens_out": 0 },
  "quality":  { "thumbs_up": 0, "thumbs_down": 0, "satisfaction_rate": 0,
                "answers_with_fabricated_citations": 0, "mean_confidence": 0 },
  "escalations": { "open": 0, "by_reason": {} },
  "by_tenant":   [ { "tenant_id": "acme", "requests": 0, … } ],
  "providers":   { "groq": { "requests": 0, "tokens_in": 0, "virtual_cost_usd": 0 } }
}
POST/admin/cache/flush?tenant_id=Drop one tenant's cached answers.

Health#

GET/healthLiveness plus an independent report per dependency.

Each dependency reports separately so a broken Redis does not mask a healthy Postgres. The endpoint reports rather than raises — a health check that can fail is one more thing to page about.

Error taxonomy#

StatusWhenShape
422Empty queryFastAPI validation detail
404Unknown tenantdetail: unknown tenant 'x' — deliberately not an empty result set
401Admin token missing or wrongThe dashboard says which side to fix
502The frontend route handler cannot reach the APIdetail naming the origin it tried
200 + error eventThe pipeline raises after the stream has begunA terminal SSE error event. The client gets a definite end state rather than a connection that just stops, which is indistinguishable from a network failure

Every environment variable and tuning knob →