Start here
Guided tour
The five cases#
What is the webhook retry limit?
Conflicting sources, resolved and disclosedWhat to watch
The sources panel populates before the first token of the answer. One of the sources carries an amber superseded in part badge. The answer should prefer the newer figure and say that the two sources disagree.
Why it works that way
This is a deliberately planted conflict. The product documentation says the retry limit is 3; the v2.4 changelog entry says 5, and it declares the conflict rather than superseding the whole page — because the changelog contradicts one fact on that page, not the page. Archiving the doc would destroy correct information to fix one stale sentence, so the conflict is pushed to generation time on purpose, where the model must prefer the newest source and flag the discrepancy. In production nobody remembers to mark the old doc; a system that only handles declared supersession handles the easy half.
What causes ERR_TIMEOUT_502?
Exact identifiers, where the keyword leg earns its placeWhat to watch
A ticket source, and an answer naming the specific cause. Compare it with make playground on the same query: the vector leg alone ranks a different error code's ticket highly, because embeddings encode meaning and two error codes mean nearly the same thing.
Why it works that way
Postgres' text-search parser splits ERR_TIMEOUT_502 into three lexemes, so a naive OR over them matches any chunk containing err. The keyword leg keeps the identifier together with a followed-by operator while OR-ing the terms the user actually typed as separate ideas. That one line took three versions to get right, and field note 1 is the whole story.
What is the capital of France?
Abstention, at zero costWhat to watch
An amber escalation banner — not red — with a ticket id, and a timing strip showing no generation stage at all. The model was never called.
Why it works that way
The confidence gate sits before generation, so a question the corpus cannot answer costs nothing but a retrieval round trip. The banner is amber because abstaining is the system working correctly; red would train a user to read correct behaviour as a fault, and that is the single most important colour decision in the application. The escalation row carries the conversation and the top ten sources with their scores, because “the right chunk was at rank 7” and “the right chunk was never retrieved” are different bugs with different fixes.
Ask the same question a second time
The cache, and instrumentation that does not lieWhat to watch
A cached badge, and total time dropping to roughly 10 ms. The cost figure for this request reads $0.000000.
Why it works that way
A cache hit records zero tokens and zero cost. Replaying the original answer's spend would make cost-per-query rise as caching improved — the dashboard would show the system getting more expensive exactly as it got cheaper, on the one metric caching exists to move. The provider and model are replayed, because “which model wrote this cached answer?” is a real debugging question.
Switch tenant, then ask 02 again
Isolation, and that it is structuralWhat to watch
The conversation clears, and the menu says so before you click. The same question now returns different private documents — or abstains, if that tenant's corpus does not cover it.
Why it works that way
Switching clears history because carrying it across would feed one tenant's answers into another's prompt as context. That is not a chunk leak, but it is a leak, and doing it silently would be the more convenient design and the less honest one. Underneath, every database read goes through a scope that owns the FROM clause; a row that surfaces outside its tenant raises rather than being quietly filtered away. The isolation page has the four layers.
The closer#
The most persuasive part of this project is not in the UI at all. Run make eval-retrieval: under a minute, no LLM calls, and it prints a scorecard comparing five retrieval strategies over the 65-case golden set.
The scorecard says hybrid retrieval lost to vector-only on this corpus — which contradicts the design document that specified hybrid retrieval, and is the most useful thing the harness produced. A system that can only confirm its own design is not being measured. The results page has every number and the caveat that bounds it.