deploy/observability/dashboards/control-plane.json (queue depth, dead letters, redeliveries, schedule-to-start, WAL size, training gauges), deploy/observability/control-plane-monitors.yaml (scrapes for the durable server, the workers, and the collector’s Prometheus exporter), and the worker chart’s export settings (dep.06): Go over OTLP/gRPC, Python over OTLP/HTTP
dep.06 (the charts these panels and monitors watch), dur.09 (telemetry.py, the Python end of the trace), obs.04 (dashboards as code; its promtool and PromQL helpers are reused); reading: obs.01 (go/otelx, the Go spans), obs.02 (scraping)
Used by
no call site (a practice): drills durable-kill9 and poison-task (ops.02, ops.03) are detected on this dashboard
Milestone
MS-durable (one trace from {ctl} train to train.step on kind)
One TrainRun is one trace: {ctl} train starts it, the durable server stores the context in the workflow’s first event, every activity task carries it, the worker hands it to Python as TRACEPARENT, and train.step spans hang under it every 50th step.
A queue-based system answers three questions on one screen: is work piling up (depth per queue), is work being given up on (dead letters), are workers dying (redeliveries, as a rate).
Schedule-to-start latency is the queue’s user-facing latency; its p95 is histogram_quantile over bucket rates grouped by le.
Python cannot be scraped: it pushes metrics over OTLP to the collector, whose Prometheus exporter is scraped instead. The Go worker speaks OTLP/gRPC on 4317, Python OTLP/HTTP on 4318.
Monitors carry release: observability, or Prometheus never sees them.
Your serving path is visible end to end (obs.01 to obs.04), but the control plane you built in Pass 8 is a black box: when a corpus build stalls, nothing says whether the queue is growing, a task is stuck in the dead-letter queue, or workers keep dying and redelivering. The training run is worse: the interesting work happens in a Python child process two hops from the CLI, and without a propagated context its spans either do not exist or start a trace of their own that nobody can connect to the TrainRun that caused them. This module makes the control plane observable: the dashboard the drills of Pass 8 are detected on, and one trace from the command you type to the training step.
stores the context in WorkflowExecutionStarted.trace_context
the StartWorkflow request’s metadata
worker (workflow task)
workflow TrainRun
the stored context, on every replay (tl.workflow.replay = true when replaying)
worker (activity task)
activity train
ActivityTask.trace_context
Python child
train.run, then train.step every 50th step, train.checkpoint
TRACEPARENT, set by the subprocess runner (dur.09)
Two links break most often: the worker not handing the activity’s context to the child (Python starts a new root trace), and the child exporting to the wrong port (its spans vanish; telemetry.py swallows export errors by design, so nothing says so).
sum by (queue, kind) (tl_durable_task_queue_depth)
is work piling up, and where
dead-letter queue size
sum by (queue) (tl_durable_dlq_size) (red at 1)
is work being given up on
redeliveries per second
sum by (queue) (rate(tl_durable_redeliveries_total[$__rate_interval]))
are workers dying or leases expiring
schedule-to-start p95
histogram_quantile(0.95, sum by (le, queue) (rate(..._bucket[$__rate_interval])))
how long tasks wait for a worker
WAL size
sum(tl_durable_wal_bytes) against walMaxBytes
how close the quota is
training loss, tokens per second
tl_train_loss, tl_train_tokens_per_second
is the run making progress
Symbol
Meaning
Type
$__rate_interval
Grafana’s rate window: at least four scrape intervals
duration
${datasource}
the dashboard’s Prometheus datasource variable
string
Every name is in otel/metrics.yaml; a panel on anything else is empty when it matters. Depth stays by (queue) because one stuck queue disappears in a total; redeliveries is a counter, so only its rate means anything.
The durable server and the workers serve /metrics on the port named health (9464), and a PodMonitor per component scrapes it. Python children live for minutes and are not addressable, so they push OTLP metrics to the collector, whose prometheus exporter (port prom-exporter, 8889) re-exposes them; a third monitor scrapes that. Every monitor carries release: observability, the label kube-prometheus-stack selects on.
Schedule-to-start p95. Over the last 5 minutes queue data saw 100 tasks with these cumulative bucket counts of tl_durable_task_schedule_to_start_seconds (rates scale them all by the same factor, so counts work by hand):
le
0.1
0.5
1
5
10
+Inf
count
40
70
88
97
100
100
The 95th task falls in the bucket (1,5]: 88 tasks are below 1 s and 97 below 5 s. Prometheus interpolates linearly inside the bucket: 1+(5−1)×97−8895−88=1+4×97≈4.11 s. Without le in the by clause the buckets are summed together and histogram_quantile returns NaN; that is test_schedule_to_start_quantile.
The trace. With TRACEPARENT=00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01 (the activity span), a 120-step run exports one train.run in trace 4bf92f...4736 whose parent is 00f067aa0ba902b7, and three train.step spans (steps 0, 50, 100) whose parent is train.run; that is test_python_span_tree_from_traceparent.