Skip to content

Observability for the control plane: workflow and activity spans, queue and DLQ panels

Moduleobs.05 · practice · ops · Pass 8 · 3 to 4 h
You builddeploy/observability/dashboards/control-plane.json (queue depth, dead letters, redeliveries, schedule-to-start, WAL size, training gauges), deploy/observability/control-plane-monitors.yaml (scrapes for the durable server, the workers, and the collector’s Prometheus exporter), and the worker chart’s export settings (dep.06): Go over OTLP/gRPC, Python over OTLP/HTTP
Contractotel/semconv.md (the span tree from {ctl} train to train.step), otel/metrics.yaml (tl.durable.*, tl.train.*), helm/observability.md (the monitor label, the Python metrics pipeline)
Testscourse/tests/obs.05/ (check runs artifacts.py: static, python, and cluster tiers; section 4)
Needsdep.06 (the charts these panels and monitors watch), dur.09 (telemetry.py, the Python end of the trace), obs.04 (dashboards as code; its promtool and PromQL helpers are reused); reading: obs.01 (go/otelx, the Go spans), obs.02 (scraping)
Used byno call site (a practice): drills durable-kill9 and poison-task (ops.02, ops.03) are detected on this dashboard
MilestoneMS-durable (one trace from {ctl} train to train.step on kind)
Optional depthTemporal: SDK metrics and schedule-to-start latency (free); OTLP specification (free); Prometheus: histograms and summaries (free)
  • One TrainRun is one trace: {ctl} train starts it, the durable server stores the context in the workflow’s first event, every activity task carries it, the worker hands it to Python as TRACEPARENT, and train.step spans hang under it every 50th step.
  • A queue-based system answers three questions on one screen: is work piling up (depth per queue), is work being given up on (dead letters), are workers dying (redeliveries, as a rate).
  • Schedule-to-start latency is the queue’s user-facing latency; its p95 is histogram_quantile over bucket rates grouped by le.
  • Python cannot be scraped: it pushes metrics over OTLP to the collector, whose Prometheus exporter is scraped instead. The Go worker speaks OTLP/gRPC on 4317, Python OTLP/HTTP on 4318.
  • Monitors carry release: observability, or Prometheus never sees them.
Terminal window
ol start obs.05 # records the start; there are no stubs
ol tests obs.05 # read the test catalog first
# write the dashboard, the monitors, and the worker chart's TL_PYTHON_OTLP_ENDPOINT (section 4)
OL_SMOKE=1 ol check obs.05 # static and python tiers
kubectl apply -f deploy/observability/control-plane-monitors.yaml
<system> train --spec specs/tiny.json # on kind: makes the trace the cluster tier looks for
ol check obs.05

Your serving path is visible end to end (obs.01 to obs.04), but the control plane you built in Pass 8 is a black box: when a corpus build stalls, nothing says whether the queue is growing, a task is stuck in the dead-letter queue, or workers keep dying and redelivering. The training run is worse: the interesting work happens in a Python child process two hops from the CLI, and without a propagated context its spans either do not exist or start a trace of their own that nobody can connect to the TrainRun that caused them. This module makes the control plane observable: the dashboard the drills of Pass 8 are detected on, and one trace from the command you type to the training step.

ProcessSpanGets its parent from
{ctl} trainthe CLI spana new root
durable serverstores the context in WorkflowExecutionStarted.trace_contextthe StartWorkflow request’s metadata
worker (workflow task)workflow TrainRunthe stored context, on every replay (tl.workflow.replay = true when replaying)
worker (activity task)activity trainActivityTask.trace_context
Python childtrain.run, then train.step every 50th step, train.checkpointTRACEPARENT, set by the subprocess runner (dur.09)

Two links break most often: the worker not handing the activity’s context to the child (Python starts a new root trace), and the child exporting to the wrong port (its spans vanish; telemetry.py swallows export errors by design, so nothing says so).

PanelQuery shapeQuestion
task queue depth by queuesum by (queue, kind) (tl_durable_task_queue_depth)is work piling up, and where
dead-letter queue sizesum by (queue) (tl_durable_dlq_size) (red at 1)is work being given up on
redeliveries per secondsum by (queue) (rate(tl_durable_redeliveries_total[$__rate_interval]))are workers dying or leases expiring
schedule-to-start p95histogram_quantile(0.95, sum by (le, queue) (rate(..._bucket[$__rate_interval])))how long tasks wait for a worker
WAL sizesum(tl_durable_wal_bytes) against walMaxByteshow close the quota is
training loss, tokens per secondtl_train_loss, tl_train_tokens_per_secondis the run making progress
SymbolMeaningType
$__rate_intervalGrafana’s rate window: at least four scrape intervalsduration
${datasource}the dashboard’s Prometheus datasource variablestring

Every name is in otel/metrics.yaml; a panel on anything else is empty when it matters. Depth stays by (queue) because one stuck queue disappears in a total; redeliveries is a counter, so only its rate means anything.

The durable server and the workers serve /metrics on the port named health (9464), and a PodMonitor per component scrapes it. Python children live for minutes and are not addressable, so they push OTLP metrics to the collector, whose prometheus exporter (port prom-exporter, 8889) re-exposes them; a third monitor scrapes that. Every monitor carries release: observability, the label kube-prometheus-stack selects on.

Schedule-to-start p95. Over the last 5 minutes queue data saw 100 tasks with these cumulative bucket counts of tl_durable_task_schedule_to_start_seconds (rates scale them all by the same factor, so counts work by hand):

le0.10.51510+Inf
count40708897100100

The 95th task falls in the bucket (1,5](1, 5]: 88 tasks are below 1 s and 97 below 5 s. Prometheus interpolates linearly inside the bucket: 1+(5−1)×95−8897−88=1+4×79≈4.111 + (5 - 1) \times \frac{95 - 88}{97 - 88} = 1 + 4 \times \frac{7}{9} \approx 4.11 s. Without le in the by clause the buckets are summed together and histogram_quantile returns NaN; that is test_schedule_to_start_quantile.

The trace. With TRACEPARENT=00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01 (the activity span), a 120-step run exports one train.run in trace 4bf92f...4736 whose parent is 00f067aa0ba902b7, and three train.step spans (steps 0, 50, 100) whose parent is train.run; that is test_python_span_tree_from_traceparent.

FileHolds
deploy/observability/dashboards/control-plane.jsonthe six panels of 2.2, uid <system>-control-plane, datasource ${datasource}
deploy/observability/control-plane-monitors.yamlPodMonitors: component durable and worker on port health, the collector on prom-exporter; each release: observability
deploy/helm/<system>-worker/values.yaml (dep.06)OTEL_EXPORTER_OTLP_ENDPOINT on 4317 for the Go worker, TL_PYTHON_OTLP_ENDPOINT on 4318 that the worker passes to Subprocess.OTLPEndpoint
TestKINDChecks
test_dashboard_is_portableunitstable uid, no id, the datasource variable on every panel
test_queries_name_contract_metricsconformanceevery metric in otel/metrics.yaml
test_queries_parseconformancepromtool parses every query
test_queue_panelsunitdepth by queue, DLQ size, redeliveries as a rate
test_schedule_to_start_quantileboundarysection 3’s quantile shape
test_storage_and_training_panelsunitWAL size and training loss panels
test_monitors_scrape_the_control_planeunitthe three monitors with the release label
test_python_exports_over_httpboundaryGo on 4317, Python on 4318
test_python_span_tree_from_traceparentconformancesection 3’s trace from your telemetry.py
test_queries_run_in_prometheusconformanceevery query evaluates on kind
test_train_trace_in_tempoconformanceone trace holds workflow TrainRun, activity train, train.run, train.step
PitfallSymptomCaught by
1. queue depth summed over queuesa stuck train queue hides behind a quiet data queuetest_queue_panels
2. the raw redeliveries counter on a grapha line that only rises; a kill loop looks like nothingtest_queue_panels
3. le dropped from the quantile’s bythe p95 panel shows NaNtest_schedule_to_start_quantile
4. Python pointed at the gRPC portPython spans and gauges vanish silentlytest_python_exports_over_http
5. the activity context not handed to the childtrain.run starts a new trace; the TrainRun trace ends at activity traintest_python_span_tree_from_traceparent, test_train_trace_in_tempo
6. a monitor without release: observabilityvalid YAML, no scrape, empty panelstest_monitors_scrape_the_control_plane
7. a panel on a metric outside the contractan empty panel during the incidenttest_queries_name_contract_metrics
DirectionModuleHow it uses this
Backdep.06the durable and worker pods these monitors scrape, and the worker’s export settings
Backdur.09telemetry.py makes the Python spans
Backobs.04dashboards as code, promtool parsing
Forwardops.02durable-kill9 is detected as a redeliveries spike on this dashboard
Forwardops.03poison-task is detected as a dead letter
Forwardops.11eventlog-disk-full shows on the WAL size panel
Your pieceProduction equivalentWhat it addsWhere to look
queue gauges from the serverTemporal’s schedule_to_start_latency and task-queue backlog metricsper-worker poll success rates, sticky-cache hit ratesTemporal “SDK metrics” reference
PodMonitorsthe OpenTelemetry Operator’s target allocatorcollectors that discover and shard scrape targets themselvesOpenTelemetry Operator docs
head sampling (every 50th step)tail sampling in the collectorkeep whole traces that errored or ran long, drop the resttailsamplingprocessor
one Grafana dashboardSLOs on queue latency with burn-rate alertsalert on schedule-to-start budget burn, not on raw depthOpenSLO, Sloth