Architecture
Caching
Two caches, cheapest and safest first#
An exact lookup is an O(1) Redis GET and cannot be wrong. A semantic lookup needs an embedding and a similarity scan and can be wrong. So the cheap safe one runs first, and the fuzzy one only on a miss.
| Layer | Key | Cost | Failure mode |
|---|---|---|---|
| Exact | tenant + the rewritten query, hashed | ~1 ms | None — it is the same question |
| Semantic | cosine similarity ≥ 0.95 against recent keys | ~15 ms | Serves an answer written for a DIFFERENT question |
The cache is keyed on the rewritten query, and sits after rewriting#
Both rewriting and the cache want to be first. Rewriting wins, and the reason is a correctness bug rather than a preference.
“What about the backoff?” means something different in a conversation about webhooks than in one about rate limits — the same four words, two correct answers. Keying on the raw query would serve one conversation's answer into another, which is a correctness failure that looks exactly like a model hallucination and would be debugged as one.
The rewritten query is standalone by construction — that is the entire property rewriting produces, and it is precisely what a cache key needs.
Two guardrails make a 0.95 threshold survivable#
Guardrail 1 — identifier-bearing queries skip the semantic cache
Embeddings encode meaning, and two error codes mean nearly the same thing.
"what causes ERR_TIMEOUT_502?"
"what causes ERR_TIMEOUT_504?"
-> embed far above 0.95 similarity
-> have completely different correct answersThis is the same weakness that justified a hybrid retrieval design in the first place. But the semantic cache is pure vector similarity — no keyword leg, no reranker, no confidence gate. It inherits the weakness with none of the mitigations that make it tolerable in retrieval.
So a broad identifier detector — error codes, version strings, HTTP status codes, ticket ids, endpoint paths — disables the fuzzy path for those queries. Deliberately broad: a false positive costs one cache miss, a false negative serves the wrong error code's answer. It is applied on write as well as read, so weakening the read-side guard later cannot detonate a stored landmine.
The exact cache still serves these queries. Only the fuzzy path, where the danger is, is off.
Guardrail 2 — abstentions are never cached
“I don't have enough information” is a statement about the corpus at one moment. Cache it and the refusal survives for an hour after someone adds the missing documentation: the system actively declines to use content it now has, and tells the user nothing exists.
Ingestion-time versioning exists to stop serving stale information. Caching a refusal would reintroduce staleness in its worst form — as a confident absence. The refusal lives in the store function, not at the call sites, so there is exactly one place that decides and a future caller cannot forget.
Active invalidation#
Every cached answer records a chunk_id → {cache keys} mapping in Redis. Re-ingesting a document collects its chunk ids and deletes precisely the answers built on them.
| Approach | Verdict |
|---|---|
| Reverse index | Precise. An answer built on five chunks dies if any one changes — conservative by design, because regenerating costs one LLM call and serving a stale answer costs trust |
| Wipe the tenant's cache on any change | Simple and correct, and re-ingestion is routine — the hit rate would spend most of its life near zero, defeating the cost argument |
| TTL only | Least code. Serves a stale answer for up to an hour after a correction ships, which is the exact failure this corpus plants conflicts to test |
Three implementation details that are the actual decision
- Chunk ids are collected before the delete. Afterwards they no longer exist, and the cache has no way to know which answers went stale.
- Archived chunks are included. An answer cached before a supersession was built on chunks now marked not-current — that answer is exactly the stale one to evict.
- Invalidation runs after the transaction commits. Inside it, a rollback would clear the cache for content that still exists, which costs a few extra LLM calls. The reverse — new content committed while stale answers survive — is much worse.
The reverse index carries a TTL of twice the cache TTL, so a late invalidation still finds something to delete. A failed invalidation is the one cache error with a real cost, so it logs at warning level rather than debug — while still never failing an ingest that has already committed.
A cache hit reports zero cost#
The cached entry carries the original provider, model and spend. Only the first two are replayed onto the new response.
Provider and model are kept, because “which model wrote this cached answer?” is a real debugging question when a cached answer turns out to be bad. The same reasoning drives what the cache entry stores at all: no retrieval result — that is a second copy of the corpus — and no per-request timings, because a cache hit took 4 ms, not the original 3,200 ms.
The event shape stays identical (meta, delta, final) so a client cannot tell the difference structurally. But the delta carries the whole answer at once: faking a typing animation for text already in hand would be adding latency to make a fast path look slow.
When Redis is unavailable#
Every Redis call is wrapped, so a cache failure degrades to a miss rather than an error. The cache can also be disabled outright with one environment variable. This matters on a free tier: Upstash's command allowance is finite and the cache spends it on every request, so exhausting the quota has to be survivable rather than fatal.
What the cache hit rate means on the dashboard, and when a change to it is a problem →