Skip to content

Migration and Cutover

How to move a customer’s production traffic from a closed API (OpenAI, Anthropic, or another provider) to an open model on a new endpoint without regressions: the endpoint swap, the template and sampling audit, shadow traffic, canary, cutover, and a rollback that has been tested. Ends with the failure-mode table to use when something goes wrong.

Parent track: Field Engineering. Follows a successful POC. Template and tokenizer depth lives in Model Loading section 7. Next: Escalation and Handoff when a failure is not yours to fix.

  • Most migrations fail on quality, not latency, and most quality failures are formatting, not model capability: chat template, stop tokens, sampling defaults, tool-call parsing.
  • An OpenAI-compatible endpoint makes the code change one line; it does not make the behavior identical. Audit every parameter.
  • Shadow traffic proves the system under real load with no user risk; canary proves it with bounded user risk. Skip neither.
  • A rollback plan is only real once it has been executed and timed.
  • Keep the incumbent’s credentials and code path live for two weeks after cutover.
  • Point an OpenAI SDK client at a local vllm serve with a small instruct model. Send the same request with and without an explicit temperature and top_p; note what changes.
  • Render a chat template locally from a model’s tokenizer_config.json and compare the token IDs to what the server reports. Find one way they could differ.
  • Walk the failure-mode table and, for each row, write the exact command or metric you would check first on your local setup.

A migration replaces a component the customer has tuned prompts against for months. Their prompts implicitly encode the incumbent’s template, defaults, and quirks. The work is to make those implicit dependencies explicit (inventory and audit), prove the replacement on real traffic without user exposure (shadow), then expose users gradually with an automatic way back (canary and rollback).

StepActionDone when
1. InventoryList every call site: model, parameters, features used (tools, JSON mode, logprobs, n, vision, seed)Spreadsheet with every call site and its parameters
2. Endpoint swapPoint the OpenAI-compatible client at the new base_url and model name behind a feature flagA single request succeeds through the customer’s real code path
3. Template and parameter auditCheck every item in section 2Every item checked and noted
4. Eval parityRun the customer’s eval on incumbent and candidate with identical inputs; diff per case, not just aggregate (04)Delta inside the agreed tolerance; every golden case passes; regressions read by a human
5. Prompt adaptationFix regressions by prompt changes first, fine-tune only if prompt changes plateauDelta inside tolerance or a scoped fine-tune proposal
6. Shadow trafficMirror a slice of production to the candidate, discard responses, log latency, errors, and outputs3+ days covering a weekly peak; no new error classes; latency within SLO
7. CanaryRoute 1%, 10%, 50% of real traffic with automatic rollback triggers on error rate and latencyEach step holds for an agreed period with product metrics flat
8. Cutover100% to candidate; keep incumbent credentials and code path liveTwo weeks clean
9. Rollback planFlag flip back to incumbent, tested during canary, with a named ownerRollback executed once in staging and timed
flowchart LR
    I["Inventory"] --> S["Endpoint swap<br/>behind a flag"]
    S --> A["Template and<br/>parameter audit"]
    A --> E["Eval parity"]
    E -->|"regressions"| PA["Prompt adaptation"]
    PA --> E
    E -->|"inside tolerance"| SH["Shadow traffic"]
    SH --> C["Canary 1% / 10% / 50%"]
    C -->|"trigger fires"| RB["Rollback to incumbent"]
    RB --> A
    C -->|"metrics flat"| CO["Cutover 100%"]
    CO -->|"two weeks clean"| D["Decommission incumbent path"]
ItemWhat goes wrongCheck
Chat templateServer applies a different template than the model was trained with; system prompt dropped or merged into the first user turnRender the template locally from the model’s tokenizer_config.json and compare to server-side token IDs
Duplicate BOSClient prepends BOS and the template adds anotherInspect the first token IDs of a rendered prompt
Stop tokensGeneration runs past the turn end, or stops early on a token that appears in contentCompare eos_token_id lists in generation_config.json versus server config
Sampling defaultsIncumbent default temperature differs from the model card’s recommended settings; top_p, top_k, repetition penalty unsetPin every sampling parameter explicitly in client code
Tool call formatModel emits tool calls in its native format; parser on the server does not match the model familyRun every tool-using eval case and assert parsed tool_calls
Structured outputJSON mode on the incumbent is not schema-constrained decoding on the new server, or vice versaValidate 100% of outputs against the schema
Reasoning tokensReasoning models emit thinking blocks the client does not strip, inflating output and costCheck for reasoning fields or tags in raw responses
TokenizerSame text, different token count: context limits, max_tokens, and cost shiftRe-count eval prompts with the new tokenizer
Context lengthServed max length is lower than the model card’sQuery the server’s model metadata; test the p99-length prompt

Key ideas:

  • Shadow at the gateway or in application code: send the request to both, return the incumbent’s answer, log the candidate’s latency, errors, finish_reason, and output. Sample outputs for human or judge comparison. Cost: the shadowed slice is paid for twice, once per provider, so size the slice deliberately.
  • Shadow data handling: production prompts now reach the new vendor. The data processing agreement, residency, and retention terms must already be signed (06).
  • Canary triggers are numbers agreed in advance: error rate above X, TTFT p95 above the SLO for Y minutes, a product metric (thumbs-down rate, escalation rate) moving more than Z. Automatic beats manual at 2 a.m.
  • Sticky routing: route by user or session, not per request, so one conversation does not switch models mid-thread.
  • Cutover is a flag value, not a deploy. The incumbent path stays deployable.
  • Rollback is rehearsed in staging during the canary and timed. Name the person who can flip it and make sure they have access outside business hours.
SymptomLikely causeFirst check
Quality drop after migrationChat template mismatch or sampling defaultsDiff rendered token IDs; pin temperature and top_p
Quality drop after quantizationOutlier-sensitive layers or KV quantization on long contextEval BF16 versus FP8 on the same cases; split long and short prompts
Outputs never stop or stop mid-sentenceWrong stop tokens or max_tokens defaultInspect finish_reason distribution
Tool calls returned as plain textTool parser does not match the model familyRaw completion versus parsed response for one case
TTFT spikes when one long prompt arrivesNo chunked prefill; long prefill blocks the batchEngine config; correlate spikes with prompt length
TTFT high at moderate loadQueueing: replica past saturation, or autoscaler too slowWaiting-queue metric; compare rate to closed-loop saturation
TPOT degrades as traffic growsBatch too large for the TPOT targetRunning sequences versus TPOT curve; cap max batch
Throughput falls, then requests restartKV cache exhausted, preemption and recomputePreemption counter, KV usage near 100%
Prefix cache hit rate near zeroDynamic content (timestamp, user ID) at the start of the promptDiff the first 100 tokens of two requests
Cost overrun on dedicatedLow batch utilization: provisioned for peak, idle off-peakGPU-hours versus tokens served by hour; autoscaling floor
Cost overrun on serverlessVolume crossed the dedicated breakeven, or output tokens grew (reasoning)Monthly tokens versus breakeven; output length trend
Cold starts after idleScale-to-zero with large weightsTime from scale event to first 200; consider a minimum of one replica
Rate-limit errors at modest loadAccount tier limits, not capacityResponse headers and tier limits
Nondeterministic outputs at temperature 0Batch-dependent numericsExpected behavior; see reproducibility
Latency fine on server, slow for usersClient-side buffering, proxy without streaming, cross-region hopsMeasure TTFT at the client and the server for one request (03)

When the first check does not explain the symptom and the remaining knobs are not yours, escalate with the 07 template.

ConceptConnected TrackApplication
Chat templates, tokenizers, generation_config.jsonModel LoadingSection 2 audit
Constrained decoding, speculative decodingLLM Systems & InferenceStructured output and latency rows
Feature flags, progressive deliveryCloud NativeCanary and rollback mechanics
Deprecation, large-scale changesSoftware CraftsmanshipRetiring the incumbent path
Alerting on SLOsObservabilityCanary triggers
CompanyHow This AppearsDifficulty
Fireworks AI / Together AI / Baseten“Migrate from OpenAI in a week” engagements; debugging template mismatchesAdvanced
OpenRouter / LiteLLM usersGateway-level provider swaps and fallbacksIntermediate
Any platform teamProgressive delivery of a critical dependencyIntermediate

Deliverable: a migration runbook for Lexa (brief) after the POC in 04 passed:

  • The inventory row for the contract-review call site, with every parameter they likely use (JSON output, max_tokens, temperature).
  • The section 2 audit filled for the chosen model, with how you would check each item.
  • Shadow plan: traffic slice, duration covering a weekday peak, what is logged, and the EU-only data path.
  • Canary steps with numeric rollback triggers, including one on JSON schema validity.
  • The rollback owner and how the rollback was tested.
  • Every canary step has an automatic trigger with a number.
  • Schema validity is checked on 100% of shadow outputs.
  • The runbook says what happens to the nightly job during cutover.
  • A peer could execute the rollback from the document alone.