Skip to content

Response cache: LRU + TTL, singleflight, tenant-scoped keys

Modulegw.06 · build · Go · Pass 7 · 4 h
You buildgo/gateway/cache/cache.go: LRU (New, Get, Put, Len, Purge), Canonicalize, KeyOf, Middleware (the cache stage with singleflight and stream replay), PurgeHandler
ContractX-TL-Cache: hit|miss on cacheable requests in course/contracts/openapi/openai-subset.v1.yaml · POST /admin/v1/cache:purge in course/contracts/openapi/admin.v1.yaml · [gateway] cache_entries, cache_ttl_s in course/contracts/config/runtime.schema.json
Testscourse/tests/go/gw_06/ (what they check: section 4); the singleflight test runs under the race detector
Needsgw.01 server skeleton (the chain, the Exchange) · gw.02 API keys (the tenant) · reading: this topic’s README, the practice drill go/02
Used bygw.07’s gateway.Deps (the Cache slot, Rev from the route epoch) and gateway.Admin (PurgeHandler)
MilestoneMS-gateway
Optional depththe Go team’s golang.org/x/sync/singleflight source (free); Megiddo and Modha, ARC: A Self-Tuning, Low Overhead Replacement Cache (FAST 2003)
  • Only deterministic requests are cached: temperature 0, or a fixed seed. Anything sampled without a seed would replay one random answer as if it were the only one (TestNotCacheableHasNoHeader).
  • The key is a hash of the tenant, the configuration revision, and the canonical request: field order, number spelling, explicit defaults, and the user field do not change it; the model, the limits, the messages, and stream do (TestHandExample, TestCanonicalDistinctions).
  • An LRU of at most cache_entries with a TTL per entry: an entry lives exactly [tput,tput+ttl)[t_{\text{put}}, t_{\text{put}} + \mathit{ttl}), and a hit makes it the most recently used (TestLRUEvictionOrder, TestTTLExpiry).
  • Singleflight: 100 identical requests arriving together make one upstream call; the others wait for it and are answered from its result (TestSingleflightOneUpstreamCall).
  • Only whole 200 answers are stored, and a stored stream is replayed byte for byte with its Content-Type (TestStreamReplayByteExact, TestErrorsAndCutStreamsNotCached).
Terminal window
ol start gw.06 # writes go/gateway/cache/cache.go with stub bodies
ol tests gw.06 # read the test catalog first
ol check gw.06 # runs the course tests with -race
ol check gw.06 --ref-deps # only if your gw.01 or gw.02 is not passing yet
ol diff gw.06 # after passing: your code against the reference

Then plug cache.Middleware(cache.New(cfg.Gateway.CacheEntries, clock), cache.Options{TTL: ..., Rev: routeEpoch}) into the Cache slot of your composition root.


The same questions reach your gateway again and again: an agent’s evaluation suite (ag.09) replays its cases at temperature 0 on every run, a docs assistant gets the same “how do I create a key?” from every new user, and the load generator (load.01) sends a fixed prompt set. Each one costs a full prefill and decode on an engine that is the most expensive thing in the system, for an answer the gateway has already seen. A response cache in front of routing turns those repeats into microseconds. It must also never do the three things that make caches dangerous: show one tenant another tenant’s answer, serve a stale answer after the model changed, or let a burst of identical misses stampede the engines.

SymbolMeaningType
NNcapacity, cache_entriesinteger, 4096 by default
ttl\mathit{ttl}lifetime of an entry, cache_ttl_sduration, 300 s by default
tpt_pwhen an entry was puttime

An LRU (least recently used) cache keeps entries in a list ordered by last use, with a map from key to list element: Get moves the hit to the front, Put inserts at the front and, when there are more than NN entries, removes from the back. Both are O(1). (Go’s container/list is the list; the map holds *list.Element.) The TTL bounds staleness: an entry is valid while now<tp+ttl\mathit{now} < t_p + \mathit{ttl}; a Get at or after that time is a miss and removes the entry. The clock is the gateway’s Clock, so tests move time with a fake.

2.2 What may be cached, and the canonical request

Section titled “2.2 What may be cached, and the canonical request”

An answer may be replayed only if asking again would give the same answer. spec/sampling.md makes that true in two cases: temperature 0 (greedy), and any temperature with a fixed seed on the same engine. Requests without either are passed through untouched, with no X-TL-Cache header.

Two requests that ask the same thing must have the same key, so the body is reduced to a canonical form:

  1. Parse it as a JSON object.
  2. Fill every omitted field that has a contract default (temperature 1, top_p 1, top_k 0, min_p 0, repetition_penalty 1, presence_penalty 0, frequency_penalty 0, n 1, stream false, logprobs null, stop null), so {} and {"top_p": 1} agree.
  3. Drop fields that do not change the answer: user.
  4. Re-encode with sorted keys: Go’s encoding/json sorts map keys, and numbers go through float64, so 0 and 0.0 agree.

stream stays in: a streamed answer is SSE bytes and a plain one is JSON, so they are different answers. The model, max_tokens, and every message stay in.

key=SHA256(tenant ∥ 0 ∥ rev ∥ 0 ∥ path ∥ 0 ∥ canonical)\mathit{key} = \mathrm{SHA256}(\mathit{tenant} \,\|\, 0 \,\|\, \mathit{rev} \,\|\, 0 \,\|\, \mathit{path} \,\|\, 0 \,\|\, \mathit{canonical})

The tenant (from gw.02’s Principal) makes the cache private per tenant: two keys of one company share answers, two companies never do, even for byte-equal requests (the prompt may contain their data; so may the answer). The revision (Options.Rev, for example the route table’s epoch from gw.05) changes when a route points a public model at a new version, so answers from the old model become unreachable at once instead of lingering for the TTL. Requests without a principal (no authn stage) are not cached at all.

On a miss, the first request for a key becomes the leader: it registers the key as in flight and goes upstream. A request for the same key that arrives while the leader runs waits for it instead of going upstream too, then is answered from the leader’s result as a hit. Without this, a popular question that just expired sends 100 identical prefills to the engines at once (a thundering herd, or cache stampede). Two details make it correct:

  • Check the cache again after taking the in-flight lock. A leader that finished between a follower’s Get and its lock has already stored the entry and left the in-flight table; without the second check, the follower would become a second leader.
  • If the leader’s answer was not cacheable (an error, a cut stream), waiters go upstream themselves rather than replaying a failure.

While the leader’s response passes to its client, a recorder keeps a copy (up to 1 MiB). It is stored only when it is a whole 200: status 200, and either valid JSON or an SSE stream that ends with data: [DONE]\n\n. A stream cut by an engine failure ends with gw.04’s error event instead and is not stored; a 503 is not stored. A hit replays the stored status, Content-Type, and body, so a streaming client gets the same SSE bytes it would have gotten from the engine (all at once). Hits set X-TL-Cache: hit, misses X-TL-Cache: miss; hits record no usage on the Exchange, so the limiter (gw.03) refunds their tokens and the meter (gw.07) records them as free.

Two requests from tenant acme:

A: {"model":"smol","messages":[{"role":"user","content":"hi"}],"temperature":0,"user":"alice"}
B: {"temperature":0.0,"top_p":1,"messages":[{"role":"user","content":"hi"}],"model":"smol"}

Canonicalize A: fill the defaults it omits (top_p 1, top_k 0, min_p 0, repetition_penalty 1, both penalties 0, n 1, stream false, logprobs and stop null), drop user, sort the keys. B already has top_p, writes 0.0 for 0, and has no user. Both become

{"frequency_penalty":0,"logprobs":null,"messages":[{"content":"hi","role":"user"}],"min_p":0,"model":"smol","n":1,"presence_penalty":0,"repetition_penalty":1,"stop":null,"stream":false,"temperature":0,"top_k":0,"top_p":1}

(inside each message the keys are sorted too). temperature is 0, so both are cacheable, and with the same tenant and revision they share one key: A is a miss that reaches the engine, B is a hit. The same body at temperature 0.7 is not cacheable; at 0.7 with "seed": 42 it is. This is TestHandExample.

package cache // import "tinyllm/gateway/cache"
type Key [32]byte
type Entry struct { Status int; ContentType string; Body []byte; Model string }
func New(capacity int, clock server.Clock) *LRU
func (c *LRU) Get(ctx context.Context, k Key) (Entry, bool)
func (c *LRU) Put(ctx context.Context, k Key, e Entry, ttl time.Duration)
func (c *LRU) Len() int
func (c *LRU) Purge(model string) int // "" = everything
type CanonicalRequest struct { Path string; JSON []byte }
func Canonicalize(path string, body []byte) (CanonicalRequest, bool /* cacheable */, error)
func KeyOf(p auth.Principal, cfgRev string, r CanonicalRequest) Key
type Options struct { TTL time.Duration; Rev func() string; MaxEntryBytes int }
func Middleware(c *LRU, o Options) server.Middleware // chat and completions only
func PurgeHandler(c *LRU) http.Handler // POST /admin/v1/cache:purge

Use only the standard library. The stage caches POST /v1/chat/completions and POST /v1/completions; embeddings and everything else pass through.

TestKINDChecksWhy it matters downstream
TestHandExampleunitsection 3: both bodies give the exact canonical JSON and one key; 0.7 without seed is not cacheable, with seed it is; no temperature means 1you and the tests agree on equality
TestCanonicalDistinctionsunitmodel, max_tokens, message text, and stream each change the canonical form; a non-object is an errordifferent questions never share an answer
TestKeyOfTenantAndRevisionunitsame tenant shares across keys; another tenant or revision does notprivacy and freshness
TestLRUEvictionOrderunitcapacity 2: put a, b; get a; put c evicts bthe hot set survives
TestTTLExpiryfaulthit at 9.999 s, miss at exactly 10 s, entry removedbounded staleness
TestPurgeunitpurge one model, then everythingPOST /admin/v1/cache:purge
TestHitAfterMissunitmiss then hit with the same body and content type, one upstream callconformance case cache.hit
TestSingleflightOneUpstreamCallfault100 concurrent identical requests: one upstream call, 100 identical bodiesno stampede after an expiry
TestTenantIsolationfaultanother tenant’s identical request is a miss that reaches the upstreamno cross-tenant leaks
TestStreamReplayByteExactunita cached stream replays the same SSE bytes with text/event-streamstreaming clients cannot tell
TestErrorsAndCutStreamsNotCachedfaulta 503 and a stream without [DONE] are not storedfailures are not replayed for 5 minutes
TestNotCacheableHasNoHeaderunita sampled request: no header, every call upstreamsampling stays random
TestPurgeHandlerunit, conformance{"purged": n} with and without a modelgw.07 mounts it
PitfallSymptomCaught by
1. a key without the tenant, or without the revisiontenant B reads tenant A’s answer; after a model rollout the old model’s answers keep comingTestTenantIsolation, TestKeyOfTenantAndRevision (mutants s01, s11)
2. hashing the raw body, skipping defaults, keeping user, or dropping streamequal questions miss; a streamed request receives a JSON bodyTestHandExample, TestCanonicalDistinctions, TestHitAfterMiss (mutants s02, s03, s10, s13, s15)
3. caching failuresone engine outage is replayed to everyone for 5 minutes; a cut stream becomes the answerTestErrorsAndCutStreamsNotCached (mutants s08, s09)
4. caching sampled requestsevery “write me a poem” gets the same poemTestNotCacheableHasNoHeader, TestHandExample (mutant s04)
5. now > expires instead of now >= expiresentries live one tick past their TTLTestTTLExpiry (mutant s05)
6. a Get that does not refresh recencythe most used entry is the first evicted (FIFO, not LRU)TestLRUEvictionOrder (mutant s06)
7. no singleflight, or no second check under the lockan expiring popular entry sends a burst of identical prefills to the enginesTestSingleflightOneUpstreamCall (mutant s07)
8. replaying without Content-TypeSSE clients do not parse the replayed streamTestStreamReplayByteExact, TestHitAfterMiss (mutant s12)
9. a purge that ignores its modelpurging one model empties the cacheTestPurge, TestPurgeHandler (mutant s14)
DirectionModuleHow it uses this
Backgw.01the Cache slot, the Exchange (tl.cache.hit), ex.Request for the body
Backgw.02Principal.Tenant scopes every key
Forwardgw.07gateway.Deps calls Middleware with Rev set to the route table’s ETag; gateway.Admin builds PurgeHandler
Forwardload.01the load report’s X-TL-Cache: hit ratio
Your pieceProduction equivalentWhat it addsWhere to look
exact-match response cachesemantic caches (GPTCache, Redis LangCache)hits for paraphrases by embedding similarity, with the false-hit risk that bringsGPTCache (free)
the gateway’s response cacheprovider prompt caching (Anthropic, OpenAI), the engine prefix cache (L8.4)reuse of the prompt’s KV, not the answer: works for sampled requests toovLLM automatic prefix caching (free)
singleflightgolang.org/x/sync/singleflight, groupcacherequest coalescing as a library; peer-to-peer cache fillinggroupcache (free)
one process’s LRUsaige’s separate response cache, prompt cache, step record, and journalidentity private by default, each cache its own contractsaige cache package