deploy/observability/dashboards/*.json: the serving dashboard (SLO status, TTFT and TPOT heatmaps, KV cache usage, per-tenant usage), as Grafana JSON in git, loaded by the stack’s Grafana from a ConfigMap
Contract
otel/metrics.yaml (the only metric names a panel may read), your obs.03 rules (recording rules and alerts), helm/observability.md (Grafana on NodePort 30300)
A dashboard is code when a fresh Grafana loads the file from git unchanged: a stable uid, no instance id, and a datasource that is a variable, not the uid of the Grafana you exported from.
Every query names only contract metrics, your recording rules, or ALERTS, and parses; the check proves both with promtool.
A heatmap of a Prometheus histogram needs bucket rates summed by (le) and the target format heatmap; the p95 line alone hides a second mode.
One row per question: are we within SLO (status), how are latencies distributed (heatmaps), is the engine full (KV blocks), who is using it (per tenant).
rate() windows follow the scrape interval: $__rate_interval, never a fixed [1m] at a 30 s scrape.
ops.01 will page you with TTFTBudgetBurnFast. The page says the budget is burning; it does not say why. The first five minutes of every incident are the same questions: which SLO, since when, is latency uniformly worse or is there a new slow mode, is the engine out of KV blocks, is one tenant flooding the gateway. Typing PromQL during an incident is slow and error-prone, and a dashboard someone clicked together in Grafana last month is gone the next time the cluster is recreated (dep.02 recreates it whenever port mappings change). Dashboards kept as JSON in your repo are reviewed, versioned, and redeployed like every other artifact.
the burn rate, slo:sli_error:ratio_rate1h{slo="ttft",slo_profile="prod"} / 0.01, thresholds at 6 and 14.4
table
which alerts fire?
ALERTS{alertstate="firing",alertname=~".+BudgetBurn(Fast|Slow)"}, instant, format table
heatmap
how are latencies distributed over time?
sum by (le) (rate(<histogram>_bucket{job="<system>-gateway"}[$__rate_interval])), format heatmap
time series
KV blocks by state, queue depth
sum by (state) (tl_engine_kv_blocks)
time series
who sends what
sum by (tenant) (rate(tl_gateway_requests_total[$__rate_interval]))
Symbol
Meaning
Type
Ci(t)
cumulative count of bucket bi (obs.02)
counter
ri=rate(Ci)
requests per second at most bi
per second
ri−ri−1
requests per second in (bi−1,bi]: one heatmap cell
per second
A heatmap cell is the difference of neighbouring cumulative rates. Grafana computes it only when the target’s format is heatmap; with the default time_series format it stacks the cumulative rates, so every cell counts all faster requests again.
rate(x[w]) uses the samples inside w; with fewer than two it returns nothing. Grafana’s $__rate_interval is at least four scrape intervals and grows with the zoom, so it is never too short. A fixed [1m] at a 30 s scrape is two samples on a good day and none after one late scrape: the panel flickers between a value and “No data”.
The stack’s Grafana runs a sidecar that loads every ConfigMap labelled grafana_dashboard: "1" as dashboards. The JSON in git is the source; the ConfigMap is generated from it (by kubectl create configmap --from-file, a Helm chart, or your Tiltfile); nothing is edited in the UI without being exported back to the file.
Two engine replicas, scraped every 10 s. Over the last minute the gateway’s TTFT buckets grew by:
le
ΔCi over 60 s
ri (per s)
cell (bi−1,bi] (per s)
0.1
300
5.0
5.0 in (0.08, 0.1] and below
0.25
540
9.0
4.0 in (0.1, 0.25]
0.5
570
9.5
0.5 in (0.25, 0.5]
0.75
570
9.5
0
1.0
594
9.9
0.4 in (0.75, 1.0]
+Inf
600
10.0
0.1 above 1.0
The heatmap column for this minute has its mass in the two lowest rows and a separate small band near one second: about 5% of requests (0.5/10) are in a second mode, here the requests that queued behind a long prefill. The p95 line misses it entirely: the 95th percentile (0.95×10=9.5 per s) lands at the 0.5 s bound. With the default format Grafana would draw 10.0, 9.9, 9.5, 9.5, 9.0, 5.0 stacked in the column, and the slow band would look like the busiest cell.
The SLO status stat next to it shows the 1 h TTFT burn: if the hour looked like this minute, e=0.5/10=0.05 of requests over 500 ms, and with β=0.01 the burn is 5: still green, one step below the orange line at 6 (the slow alert’s factor); red starts at 14.4 (the fast one’s).
heatmap: TTFT; heatmap: TPOT; time series: TTFT and TPOT p95
Engine capacity
KV cache blocks by state (stacked); queue depth and running sequences per pod
Tenants
requests per second by tenant; 5xx per second by tenant
Panel and target datasource: {"type": "prometheus", "uid": "${datasource}"} with a datasource variable of type datasource, query prometheus, or the stack’s provisioned uid prometheus. Before checking, the course replaces $__rate_interval, $__interval, and $__range with 5m and every other variable with .*.