Skip to content

Dashboards as code for the serving path

Moduleobs.04 · practice · ops · Pass 7 · 3 to 4 h
You builddeploy/observability/dashboards/*.json: the serving dashboard (SLO status, TTFT and TPOT heatmaps, KV cache usage, per-tenant usage), as Grafana JSON in git, loaded by the stack’s Grafana from a ConfigMap
Contractotel/metrics.yaml (the only metric names a panel may read), your obs.03 rules (recording rules and alerts), helm/observability.md (Grafana on NodePort 30300)
Testscourse/tests/obs.04/ (check runs artifacts.py; section 4)
Needsobs.02 (the series), obs.03 (the burn-rate recording rules and the six alerts)
Used byops.01 and every later drill start from this dashboard; ops.09 (noisy neighbor) from its tenant row
MilestoneMS-prod (Grafana during a load run)
Optional depthGrafana dashboard JSON model (free), Heatmaps of Prometheus histograms (free), The RED method and The USE method (free articles)
  • A dashboard is code when a fresh Grafana loads the file from git unchanged: a stable uid, no instance id, and a datasource that is a variable, not the uid of the Grafana you exported from.
  • Every query names only contract metrics, your recording rules, or ALERTS, and parses; the check proves both with promtool.
  • A heatmap of a Prometheus histogram needs bucket rates summed by (le) and the target format heatmap; the p95 line alone hides a second mode.
  • One row per question: are we within SLO (status), how are latencies distributed (heatmaps), is the engine full (KV blocks), who is using it (per tenant).
  • rate() windows follow the scrape interval: $__rate_interval, never a fixed [1m] at a 30 s scrape.
Terminal window
ol start obs.04 # records the start; there are no stubs
ol tests obs.04 # read the test catalog first
# build the dashboard in Grafana (http://127.0.0.1:30300), export JSON, clean it (section 5), then:
OL_SMOKE=1 ol check obs.04 # the static tier
kubectl -n observability create configmap forge-dashboards --from-file=deploy/observability/dashboards \
--dry-run=client -o yaml | kubectl label --local -f - grafana_dashboard=1 -o yaml | kubectl apply -f -
ol check obs.04 # adds: queries run in Prometheus, the ConfigMap is there

ops.01 will page you with TTFTBudgetBurnFast. The page says the budget is burning; it does not say why. The first five minutes of every incident are the same questions: which SLO, since when, is latency uniformly worse or is there a new slow mode, is the engine out of KV blocks, is one tenant flooding the gateway. Typing PromQL during an incident is slow and error-prone, and a dashboard someone clicked together in Grafana last month is gone the next time the cluster is recreated (dep.02 recreates it whenever port mappings change). Dashboards kept as JSON in your repo are reviewed, versioned, and redeployed like every other artifact.

FieldMeaningRule
uidthe dashboard’s identity in URLs and provisioningstable, set by you, at most 40 of [A-Za-z0-9_-]
idthe database row id in one Grafananull in git: another Grafana assigns its own
panels[]each panel: type, title, gridPos, datasource, targets[]type: row panels group the others
targets[].exprthe PromQL a panel runscontract metrics, recording rules, ALERTS only
templating.list[]variables, e.g. datasource of type datasource, query prometheuspanels refer to ${datasource}
PanelQuestionQuery shape
statis this SLO burning now?the burn rate, slo:sli_error:ratio_rate1h{slo="ttft",slo_profile="prod"} / 0.01, thresholds at 6 and 14.4
tablewhich alerts fire?ALERTS{alertstate="firing",alertname=~".+BudgetBurn(Fast|Slow)"}, instant, format table
heatmaphow are latencies distributed over time?sum by (le) (rate(<histogram>_bucket{job="<system>-gateway"}[$__rate_interval])), format heatmap
time seriesKV blocks by state, queue depthsum by (state) (tl_engine_kv_blocks)
time serieswho sends whatsum by (tenant) (rate(tl_gateway_requests_total[$__rate_interval]))
SymbolMeaningType
Ci(t)C_i(t)cumulative count of bucket bib_i (obs.02)counter
ri=rate(Ci)r_i = \text{rate}(C_i)requests per second at most bib_iper second
ri−ri−1r_i - r_{i-1}requests per second in (bi−1,bi](b_{i-1}, b_i]: one heatmap cellper second

A heatmap cell is the difference of neighbouring cumulative rates. Grafana computes it only when the target’s format is heatmap; with the default time_series format it stacks the cumulative rates, so every cell counts all faster requests again.

rate(x[w]) uses the samples inside ww; with fewer than two it returns nothing. Grafana’s $__rate_interval is at least four scrape intervals and grows with the zoom, so it is never too short. A fixed [1m] at a 30 s scrape is two samples on a good day and none after one late scrape: the panel flickers between a value and “No data”.

The stack’s Grafana runs a sidecar that loads every ConfigMap labelled grafana_dashboard: "1" as dashboards. The JSON in git is the source; the ConfigMap is generated from it (by kubectl create configmap --from-file, a Helm chart, or your Tiltfile); nothing is edited in the UI without being exported back to the file.

Two engine replicas, scraped every 10 s. Over the last minute the gateway’s TTFT buckets grew by:

leΔCi\Delta C_i over 60 srir_i (per s)cell (bi−1,bi](b_{i-1}, b_i] (per s)
0.13005.05.0 in (0.08, 0.1] and below
0.255409.04.0 in (0.1, 0.25]
0.55709.50.5 in (0.25, 0.5]
0.755709.50
1.05949.90.4 in (0.75, 1.0]
+Inf60010.00.1 above 1.0

The heatmap column for this minute has its mass in the two lowest rows and a separate small band near one second: about 5% of requests (0.5/100.5 / 10) are in a second mode, here the requests that queued behind a long prefill. The p95 line misses it entirely: the 95th percentile (0.95×10=9.50.95 \times 10 = 9.5 per s) lands at the 0.5 s bound. With the default format Grafana would draw 10.0, 9.9, 9.5, 9.5, 9.0, 5.0 stacked in the column, and the slow band would look like the busiest cell.

The SLO status stat next to it shows the 1 h TTFT burn: if the hour looked like this minute, e=0.5/10=0.05e = 0.5 / 10 = 0.05 of requests over 500 ms, and with β=0.01\beta = 0.01 the burn is 55: still green, one step below the orange line at 6 (the slow alert’s factor); red starts at 14.4 (the fast one’s).

deploy/observability/dashboards/serving.json (the reference’s rows):

RowPanels
SLOsstat: TTFT, TPOT, and availability burn rate (1 h); table: firing *BudgetBurn* alerts
Latencyheatmap: TTFT; heatmap: TPOT; time series: TTFT and TPOT p95
Engine capacityKV cache blocks by state (stacked); queue depth and running sequences per pod
Tenantsrequests per second by tenant; 5xx per second by tenant

Panel and target datasource: {"type": "prometheus", "uid": "${datasource}"} with a datasource variable of type datasource, query prometheus, or the stack’s provisioned uid prometheus. Before checking, the course replaces $__rate_interval, $__interval, and $__range with 5m and every other variable with .*.

TestKINDChecksWhy it matters downstream
test_dashboards_are_portable_grafana_jsonunitvalid JSON, unique stable uid, id null, title, schemaVersion, panelsthe file loads in any Grafana
test_panels_use_a_portable_prometheus_datasourceunitevery panel and target uses a datasource variable or uid prometheusno “datasource not found” after a rebuild
test_every_query_names_contract_metricsconformancemetric names are in metrics.yaml, your recording rules, or ALERTSno empty panel that looks like “no traffic”
test_every_query_parsesconformancepromtool parses every querysyntax errors show up now, not mid-incident
test_slo_status_panels_cover_the_three_slosunita status panel (stat, gauge, bar gauge, table, state timeline) per SLOthe first question of every page
test_latency_heatmapsunitTTFT and TPOT heatmaps: bucket rates by (le), format heatmapsection 3’s second mode
test_kv_cache_usage_panelunita panel reads tl_engine_kv_blockscapacity, leaks, preemption pressure
test_per_tenant_usage_panelunittl_gateway_requests_total by tenantops.09, noisy neighbors
test_rate_windows_survive_the_scrape_intervalboundaryevery rate/irate/increase window is $__rate_interval or at least 2 mno flickering panels
test_queries_run_in_prometheusconformanceon kind: each query runs in [deploy].prometheus without errormany-to-many and runtime errors
test_dashboards_provisioned_for_grafanaconformanceon kind: a ConfigMap labelled grafana_dashboard=1 holds each uidthe dashboard is on screen during a drill
PitfallSymptomCaught by
1. Exported with "id": 17 and a generated uida second import overwrites or duplicates; links break on every re-exporttest_dashboards_are_portable_grafana_json
2. The exporting Grafana’s datasource uid in every panel“datasource not found” on a fresh clustertest_panels_use_a_portable_prometheus_datasource
3. A metric name from a blog post or a typo (ttft_seconds_bucket)an empty panel that reads as zero traffictest_every_query_names_contract_metrics
4. A heatmap without by (le) or with the time series formatthe busiest-looking cell is the slow tailtest_latency_heatmaps
5. Usage summed over all tenantsone tenant’s flood looks like organic growthtest_per_tenant_usage_panel
DirectionModuleHow it uses this
Backobs.02every series a panel reads, under its job names
Backobs.03slo:sli_error:* recording rules and the six alerts in the status row
Forwardops.01the drill’s diagnosis starts at the TTFT heatmap and the queue depth per pod
Forwardops.09the tenant row shows the noisy neighbor
Forwardobs.05the control-plane dashboard (queues, DLQ) follows the same rules
Your pieceProduction equivalentWhat it addsWhere to look
hand-edited JSONGrafonnet (Jsonnet), the Grafana Foundation SDK, Grafana’s grafanactldashboards generated from code, shared panel librariesGrafonnet
a ConfigMap per folderGrafana provisioning from Git, or the Grafana Operatorsync, folders, permissions as codeGrafana Operator
one serving dashboardRED (rate, errors, duration) per service, USE (utilization, saturation, errors) per resourcea standard layout every team reads the same wayThe RED Method (Wilkie), The USE Method (Gregg)
burn-rate statsSLO dashboards from Sloth or Pyrrabudget remaining over the period, burn historyPyrra