Skip to content

Escalation and Handoff

How to hand a problem to the people who can fix it so they act on the first read: the escalation report for the performance team, the evidence each counterpart needs, the incident call, and how to relay research findings back to a customer.

Parent track: Field Engineering. Uses the benchmark evidence from 03 and the failure-mode table from 05.

  • Escalate when you have done the work you can do, and bring evidence, not a narrative.
  • A handoff is accepted when the receiver can act without calling you.
  • The “what I ruled out” section separates an escalation from a forwarded complaint.
  • During an incident the customer needs an update cadence more than a root cause. Commit to a time and keep it.
  • When taking over work, ask for the same artifacts you would hand over. If they do not exist, writing them is your first deliverable.
  • Write an escalation for a problem you can reproduce locally (for example, loadgen.py --mock with --mock-max-running set low so TTFT explodes). Hand it to a peer who has not seen your setup and time how long until they ask a question.
  • Read one public postmortem per day for a week and rewrite its summary in the incident update format below.
  • Practice relaying a technical finding at two levels: one sentence for the customer’s VP, one paragraph for their engineer.

Specialists (performance engineers, MLEs, researchers) are the scarcest people in an inference company, and every round trip with them costs a day. The field engineer’s leverage is to arrive with the environment, the traffic shape, the measurement, and the hypotheses already ruled out, so the specialist starts at diagnosis instead of interrogation. The same discipline runs the other way: findings from research only help a customer when they are translated into a decision the customer can make.

# Escalation: <one-line symptom with a number>
e.g. "TPOT p95 doubles to 80 ms above 40 req/s on <model> dedicated, customer <X>"
Impact: <customer, users affected, SLO breached, since when>
Severity: <sev level and why>
## Environment
- Engine + version/commit: <...>
- Launch flags (full, verbatim): <...>
- Model + revision + quantization: <...>
- GPU type, count, TP/PP/EP, region, node IDs: <...>
## Traffic shape
- Rate: <req/s avg/peak>, arrival pattern: <steady/bursty>
- Input tokens p50/p95/max, output tokens p50/p95/max
- Prefix sharing: <length, hit rate if known>
- Streaming: <yes/no>; features: <tools, JSON schema, LoRA>
## Observed vs expected
| Metric | Expected | Observed | Source |
## Evidence
- Benchmark command and raw output: <link>
- Engine metrics over the window (queue depth, running/waiting seqs, KV usage, preemptions): <link>
- Trace or profile: <link>
- Minimal repro: <script that reproduces with synthetic or sanitized data>
## What I ruled out
- <hypothesis> -> <how ruled out>
## Ask
<what you need: root cause, config recommendation, patch, ETA>

The “expected” column needs a source: the benchmark report (03), the hypothesis math (02), or a previous run. “It feels slow” is not an expected value.

ToWhenBring
MTS / performanceLatency or throughput misses SLO after config tuning you can do yourselfEngine version and commit, full launch flags, GPU type and count, traffic shape (QPS, input/output token histograms, prefix share), benchmark command and raw output, a trace or profile, minimal repro script
MLEQuality gap after migration or fine-tune not convergingEval set (or a sanitized sample), metric definition, baseline score versus current score, prompts with chat template applied, sampling params, training config and loss curves
ResearchA customer problem no current recipe solves (a new architecture, a domain where speculators fail, a quantization that breaks a capability)The pattern across customers, the measured gap, why the existing recipes were ruled out, the business value of solving it
ProductA missing feature or a recurring workaround blocks dealsAccounts affected, the workaround and its cost, revenue at stake, the smallest change that removes the blocker
Solutions architectA pattern repeats across three or more accountsThe three accounts, the common integration, what you built each time
Account executiveScope or commercial change: new workload, capacity request, SLO changeRevised deployment hypothesis with GPU count and cost per 1M tokens, risk list, timeline

The reverse direction matters too. When you take over from an MLE or MTS, ask for the same artifacts. If they do not exist, writing them is your first deliverable.

Before escalating to performance, walk the failure-mode table and record each first check in “what I ruled out”. Most escalations that bounce back fail one of these: no engine version, no traffic shape, a closed-loop benchmark offered as SLO evidence, or prefix-cache hits in the repro.

Agenda (30 min, can be called within the hour)

MinItem
0-5Impact: what users see, since when, how many requests
5-15Timeline and what changed (their deploy, our deploy, traffic shift)
15-25Current mitigation and next diagnostic step, with owners
25-30Next update time; who is on point on each side

Key ideas:

  • One voice to the customer. Name the person who sends updates; specialists work the problem, they do not field status questions.
  • Cadence over cause. Commit to “next update at 14:30” and keep it, even if the update is “no change”.
  • Mitigate first. Roll back, shift traffic, add replicas, or fall back to the incumbent path (05) before root cause is known.
  • Blameless write-up after. Timeline, impact, root cause, contributing factors, action items with owners (Postmortem Culture).

Update format:

[<time, timezone>] <customer> / <symptom> / status: investigating | mitigated | resolved
Impact now: <what users see, % of requests>
Since last update: <what was done, what was learned>
Next step: <action, owner>
Next update: <time>

Research and performance findings arrive as mechanisms (“the speculator’s acceptance rate drops on legal text because the draft was trained on chat data”). The customer needs a decision.

LayerExampleAudience
FindingSpeculative decoding acceptance rate is 0.35 on their traffic versus 0.7 on chat benchmarksTheir MLE, in a written note
ConsequenceTPOT gain at low load is about 10%, not the 2x on the vendor blogTheir engineering lead
DecisionKeep speculative decoding off at peak; test a draft fine-tuned on their domain next quarter, or accept current TPOTTheir decision maker
What changes in the planHypothesis row updated; no change to cost; new risk loggedEveryone, in the next notes

Rules:

  • Translate, do not forward. A research thread pasted into a customer channel leaks internal context and confuses the reader.
  • State what was measured and what was not; carry the “not confirmed” marker through.
  • Ask research before sharing anything unpublished, and never share another customer’s data or name as the evidence.
  • Close the loop internally: tell research what the customer decided, so the pattern counts toward their priorities.
ConceptConnected TrackApplication
Engine metrics: queue depth, KV usage, preemptionsServing & LoadEvidence section
Traces, dashboards, alertingObservabilityEngine metrics and profiles
Postmortems, lessons learnedSoftware Craftsmanship: Lessons from PracticeBlameless write-ups
Eval artifacts for MLE handoffsLLM EvaluationQuality escalations
CompanyHow This AppearsDifficulty
Fireworks AI / Together AI / BasetenFDE-to-performance-team escalations; “how would you escalate this?” interview promptsAdvanced
Google / any SRE organizationIncident command, update cadence, blameless postmortemsIntermediate
PalantirField-to-product feedback loops from forward deployed teamsIntermediate

Scenario: Lexa (brief) is two weeks past cutover (05). Every weekday around 9:30 CET, TTFT p95 for contract review rises from about 1.2 s to 6 s for twenty minutes, then recovers. Error rate is flat. Their nightly job finished late twice this week.

Deliverable:

  1. An escalation report to the performance team using the template, with plausible values from your earlier documents and at least four hypotheses in “what I ruled out” (each with the check you ran, drawn from the failure-mode table).
  2. The first two incident updates to Lexa in the update format.
  3. A relay note for Lexa’s VP of Engineering, assuming performance finds that autoscaling reacts to GPU utilization instead of queue depth and the nightly job overran into the morning.
  • A performance engineer could start diagnosis from the report without a question (test with a peer).
  • The report links the overrunning nightly job to the morning symptom as a hypothesis with evidence, not a conclusion.
  • Each incident update has a next-update time.
  • The relay note ends in a decision Lexa can make, with its cost.