Skip to content

Start here

Guided tour

Five queries, in this order. Each one turns on a defence the previous one did not need, so running them out of sequence makes several of them look like nothing is happening. About three minutes end to end.

The five cases#

01

What is the webhook retry limit?

Conflicting sources, resolved and disclosed

What to watch

The sources panel populates before the first token of the answer. One of the sources carries an amber superseded in part badge. The answer should prefer the newer figure and say that the two sources disagree.

Why it works that way

This is a deliberately planted conflict. The product documentation says the retry limit is 3; the v2.4 changelog entry says 5, and it declares the conflict rather than superseding the whole page — because the changelog contradicts one fact on that page, not the page. Archiving the doc would destroy correct information to fix one stale sentence, so the conflict is pushed to generation time on purpose, where the model must prefer the newest source and flag the discrepancy. In production nobody remembers to mark the old doc; a system that only handles declared supersession handles the easy half.

02

What causes ERR_TIMEOUT_502?

Exact identifiers, where the keyword leg earns its place

What to watch

A ticket source, and an answer naming the specific cause. Compare it with make playground on the same query: the vector leg alone ranks a different error code's ticket highly, because embeddings encode meaning and two error codes mean nearly the same thing.

Why it works that way

Postgres' text-search parser splits ERR_TIMEOUT_502 into three lexemes, so a naive OR over them matches any chunk containing err. The keyword leg keeps the identifier together with a followed-by operator while OR-ing the terms the user actually typed as separate ideas. That one line took three versions to get right, and field note 1 is the whole story.

03

What is the capital of France?

Abstention, at zero cost

What to watch

An amber escalation banner — not red — with a ticket id, and a timing strip showing no generation stage at all. The model was never called.

Why it works that way

The confidence gate sits before generation, so a question the corpus cannot answer costs nothing but a retrieval round trip. The banner is amber because abstaining is the system working correctly; red would train a user to read correct behaviour as a fault, and that is the single most important colour decision in the application. The escalation row carries the conversation and the top ten sources with their scores, because “the right chunk was at rank 7” and “the right chunk was never retrieved” are different bugs with different fixes.

04

Ask the same question a second time

The cache, and instrumentation that does not lie

What to watch

A cached badge, and total time dropping to roughly 10 ms. The cost figure for this request reads $0.000000.

Why it works that way

A cache hit records zero tokens and zero cost. Replaying the original answer's spend would make cost-per-query rise as caching improved — the dashboard would show the system getting more expensive exactly as it got cheaper, on the one metric caching exists to move. The provider and model are replayed, because “which model wrote this cached answer?” is a real debugging question.

05

Switch tenant, then ask 02 again

Isolation, and that it is structural

What to watch

The conversation clears, and the menu says so before you click. The same question now returns different private documents — or abstains, if that tenant's corpus does not cover it.

Why it works that way

Switching clears history because carrying it across would feed one tenant's answers into another's prompt as context. That is not a chunk leak, but it is a leak, and doing it silently would be the more convenient design and the less honest one. Underneath, every database read goes through a scope that owns the FROM clause; a row that surfaces outside its tenant raises rather than being quietly filtered away. The isolation page has the four layers.

The closer#

The most persuasive part of this project is not in the UI at all. Run make eval-retrieval: under a minute, no LLM calls, and it prints a scorecard comparing five retrieval strategies over the 65-case golden set.

The scorecard says hybrid retrieval lost to vector-only on this corpus — which contradicts the design document that specified hybrid retrieval, and is the most useful thing the harness produced. A system that can only confirm its own design is not being measured. The results page has every number and the caveat that bounds it.