Skip to content

Model Routing & Cascades

  • Routing is three decisions, not one. Which model answers, which endpoint serves it, and what happens after a failure. Products blur them; keep them apart when you evaluate.
  • The oracle proves opportunity, not a router. An oracle picks the best model after seeing every outcome. A real router chooses before. The distance between them is the whole engineering problem.
  • The common failure is model recall, not difficulty estimation. Routers tend to fail because they do not reach for the model that would have succeeded, even when it is in the pool.
  • Most of the value is in two or three complementary models. Coverage of different capabilities beats a long candidate list.
  • A model switch can cost more than it saves. Changing models abandons the prompt cache. A router that ignores cache state optimises the sticker price and raises the bill.
  • Escalate on category or on a verifier, not on model confidence alone. Confidence is produced by the thing you are trying to check, and untrusted input can move it.
  • The router is a small share of the cost and decides the rest. That makes precision the adoption metric and the labelled set the durable asset.
  • A decision is not a generation. Routing, gating, and scoring have a closed answer set. A model that returns a typed choice with a probability fits that job better than one that writes text you then parse.
  • The pool is filling with specialised models. Open bases post-trained on a verifiable reward now sit beside general frontier models, and they are often the cheap tier worth routing to.
  • Take an eval set you already have and run three models over it. Build the success matrix (task by model), then compute best single model, oracle, and best pair by hand. The gap between the first two is your opportunity.
  • Write the simplest router that could work: a rule on one input feature. Measure how often it selects the non-default model when that model was the only one that would have succeeded. That number is model recall.
  • Price one long session twice: once staying on a single model with warm cache, once switching models halfway. Use real cached-input rates.
  • Write down your escalation policy before you build anything, including what happens when the router itself errors.

No single model is best at everything, and prices differ by more than an order of magnitude. If you could always send each request to the cheapest model that would succeed, you would get higher accuracy than the best model at a fraction of its cost. That statement is true and nearly useless, because it describes an oracle. Routing is the discipline of approximating the oracle with information available before the answer exists, and of knowing how much of the gap you actually closed.

DecisionQuestionTypical owner
Model selectionWhich model should answer this request?A router: rules, a classifier, or a learned scorer
Provider selectionWhich endpoint serves the chosen model?A gateway, on price, availability, or latency
FallbackWhat runs after an error or timeout?An ordered list supplied by the caller

Pinning a model name does not pin the provider, the quantization, or the cache. Using a gateway does not mean a learned router chose the model. When someone says “we use routing,” ask which of the three they mean.

For where these decisions sit in a full serving stack, and how gateways, orchestrators, and inference engines differ, see LLM Serving Platforms. This track covers the decision itself.

FamilyDecidesSignalExtra costMain weakness
Rules and categoriesBefore generationDeterministic features: file paths, tool names, request typeNoneRules drift and do not generalise
Predictive routerBefore generationA model scores the promptOne small inferenceMisjudges prompts unlike its training data
CascadeAfter generationA scorer judges the cheap answer, then escalatesThe cheap attempt is paid even when discardedLatency, and a scorer that can be wrong
Verifier-gatedAfter generationA deterministic check: tests pass, query executes, schema validatesThe cheap attempt plus the checkOnly works where a verifier exists
Session-level routerPer task or per user turnTask features plus cache stateOne decision per turnEarly choices shape every later step

Predictive routing: RouteLLM. A router estimates how much a strong model would beat a weak one on this prompt, and a threshold turns that score into a choice. The project trained four routers on Chatbot Arena preference data: similarity-weighted ranking, matrix factorization, a BERT classifier, and a causal LLM classifier. Reported results against a GPT-4 baseline at 95% of its quality were cost reductions of over 85% on MT Bench, 45% on MMLU, and 35% on GSM8K. The spread is the lesson: savings depend on the workload. Routers trained only on Arena data performed poorly on MMLU until about 1,500 in-domain examples were added, under 2% of the training set.

Cascades: FrugalGPT. Call models in order from cheap to expensive and stop when a scoring function accepts the answer. The paper reports matching GPT-4 with up to 98% cost reduction, or 4% better accuracy at equal cost, on its evaluated tasks. A cascade pays for every rejected attempt, so it wins only when the cheap model is accepted most of the time.

A retry loop is a cascade in time. Returning an error to the same model and retrying, as in the Grounded SQL Agent, escalates effort. A cascade escalates tier. Both need a hard budget.

An oracle router assigns each task to the model with the best measured outcome, breaking ties on cost. It needs the answer key, so it is an upper bound and nothing more.

Fireworks published an oracle analysis on DeepSWE v1.1: 113 agentic software-engineering tasks, 18 models, four rollouts per pair.

PolicySuccessCost per task
Best fixed model (GPT-6 Astra)74.1%$6.52
Claude Opus 5, fixed73.8%$11.84
Oracle over open-weight models only90.3%$1.45
Oracle over all 18 models97.6%$1.88

Portfolio results from the same analysis: the best pair added 13.1 points over the best single model, the best trio reached 91.2%, and going from three models to eighteen added only 6.4 more points.

Three things to take from it:

  1. Complementarity is real. Models fail on different tasks, so the union of a pool covers far more than its best member.
  2. Curation beats count. Most of the gain arrives by the third model.
  3. None of this measures a router. The analysis modelled the simplest policy, one model chosen at the start of a task and held. Whether a deployed router captures the opportunity is a separate measurement.

The same post reports production numbers: 2,334 internal coding sessions over four weeks at $7.42 per session, against $15.81 for a fixed Claude Opus 5, a 53% reduction. That is a cost comparison. Before treating any routing claim as a quality claim, ask for task success on the same sessions.

LLMRouterBench re-evaluated the field under one framework: over 400K instances, 21 datasets, 33 models, and 10 routing baselines. Its findings are the counterweight to every oracle chart:

  • Many routing methods perform about the same once evaluated identically.
  • Several recent approaches, including commercial routers, fail to reliably beat the best single model.
  • A large gap to the oracle persists, driven mainly by model-recall failures: the right model was in the pool and the router did not choose it.
  • The embedding backbone matters little. Larger pools show diminishing returns against careful curation.

Baselines every router must beat: the best single model, a random split at the same cost, and a rule a person could write in ten minutes. If the router does not beat all three on your traffic, it is complexity without a return.

Difficulty is the wrong target. Predicting that a prompt is hard does not tell you which model handles it. A router needs to learn what each model is good at, which is why in-domain labelled outcomes matter more than general preference data.

UnitExampleTrade-off
Per requestClassify each API callMost flexible; breaks cache and consistency in multi-step work
Per user turnOne model handles the turn and its tool callsKeeps a tool loop coherent; cannot correct mid-turn
Per task or sessionPick once, holdBest cache reuse; a wrong first choice is expensive

In agent workloads, early decisions change later inputs. A model that reads the wrong file produces a different context for every following step, so per-step accuracy numbers do not compose. Evaluate complete trajectories.

A router that scores only the last message has a specific blind spot: when that message is “continue” or a tool result, it carries almost no signal about the task.

Prompt caching bills previously seen prefix tokens at a reduced rate. Caches are per model, and usually per provider. Switching models discards the cache and pays full prefill on the new one.

Illustrative arithmetic, with made-up round rates:

Option for the next turn, 100K tokens of contextInput rateInput cost
Stay on the strong model, cache hit$0.50 per 1M (cached)$0.05
Switch to a cheap model, cache miss$0.50 per 1M (uncached)$0.05
Switch back to the strong model later, cache expired$5.00 per 1M (uncached)$0.50

The cheap model saved nothing on input for that turn, and the round trip cost ten times the turn it tried to optimise. A router has to price the state it abandons, not only the model it selects. This is why session-level routers hold a choice across a turn, and why “cheapest model per request” is often the most expensive policy for long contexts.

Record in every evaluation: cache hits, cache writes, the resolved model, and the resolved provider.

Escalation is the rule for leaving the cheap path. There are four ways to trigger it, and they fail differently.

TriggerHow it worksFails when
ConfidenceEscalate when the model or a scorer reports low confidenceConfidence is miscalibrated, or the input manipulates it
CategoryNamed categories always escalate, whatever the model thinksThe category list is incomplete
VerifierA deterministic check rejects the outputNo verifier exists for the task
Human overrideA person corrects the routeVolume exceeds attention

Category beats confidence on untrusted input. If the content being routed can contain instructions, a confidence score is attacker-influenced. A rule such as “security and schema changes always get the strong reviewer” cannot be argued with, because no model evaluates it. This is the same principle as the deterministic gate in the Grounded SQL Agent: the control sits outside the model.

Decide the failure direction explicitly.

SystemRouter errors, thenWhy
Review or analysis fan-outFail open: run everythingWorst case is a known cost, never silent loss of coverage
Anything that executes or writesFail closed: refuseWorst case must not be an unreviewed action

Treat the escalation rate as a metric. A cost model that assumes 20% of calls escalate is wrong the day the rate becomes 40%. Alert on it, and cap it with a budget.

An override is also a label. When a person corrects a route, record the pair. The escape hatch and the training data are the same mechanism.

8. Economics: the Small Decision That Steers the Bill

Section titled “8. Economics: the Small Decision That Steers the Bill”

Work from tokens, not from a price list. Every per-request figure is tokens times a rate, so every figure can be re-derived when a rate changes.

Worked example: a system with six specialised subagents, of which two are relevant to a typical request. Rates are illustrative.

CallTokens in / outRate per 1M in / outCost
Frontier subagent15K / 1K$2.50 / $15.00$0.0525
Small tuned subagent9K / 1K$0.20 flat$0.0020
Prompted classifier, small model19K / 0.1K$0.10 flat$0.0019
Tuned classifier, small model2K / 0.1K$0.10 flat$0.0002
ConfigurationPer requestDerivation
Run all six on the frontier model$0.3156 × $0.0525
Prompt-routed, two frontier subagents$0.107$0.0019 + 2 × $0.0525
Tuned router, cascade, 20% escalate$0.024$0.0002 + 2 × (0.8 × $0.0020 + 0.2 × $0.0525)

Three independent levers produced that range:

  1. Routing: which calls run at all.
  2. Tier: which model runs them.
  3. Context: how many tokens each call carries. Instructions and examples re-sent on every call are what tuning moves into weights.

The router is about 1% of the final figure and determines the other 99%.

Precision pays twice. A wrong route costs the call and produces a wrong output that someone has to read. At 20 outputs a week, 90% precision is two wrong ones every week, and 99% is one every five weeks. That is the difference between a tool people read and a tool people mute.

Coverage and precision have to land together. Fixing coverage alone means more calls, a larger bill, and more wrong outputs at once.

Check how a tuned model is served before you trust a per-token number. Some platforms serve adapters only on dedicated deployments billed by GPU time. At $7 per GPU-hour, that is about $5,040 a month whether or not traffic arrives, and it changes the break-even completely.

Two choices are independent: how much of the model changes, and what signal it trains against.

How much changesWhat it isUse when
LoRABase frozen, small low-rank adapters trained; artifact in megabytesNarrow specialisation
Full-parameterEvery weight updated; a full checkpointNew capability or domain
SignalNeedsExample for a router
SFTInput and correct output pairsRequest and the route a person chose
DPOChosen and rejected pairsThe route that was overridden against its replacement
RFTA scorer that returns a reward from 0 to 1Did the downstream task succeed on this route?

Climb in order, and only when the harness shows the rung below has stopped paying:

  1. Prompted classifier. No weights change, no training data required. It sets the baseline.
  2. SFT on historical labels. Available as soon as labelled routes exist.
  3. Preference and reward tuning. DPO on overrides, RFT on outcomes. These need live traffic, so they cannot be pulled forward.

RFT in one paragraph. You supply prompts and an evaluator, and training increases the score. Gold completions are optional, which makes RFT the fit when a task is verifiable and you lack reference answers. The evaluator can be rules, tests, a judge model, or a mix. The risk is reward hacking: the model learns to satisfy the evaluator instead of the task. Inspect rollouts, and stop when quality regresses or the evaluator saturates. A good eval harness is already most of a reward function.

Be honest about label quality. Labels that come from the subset of users who bother to label skew toward trained users, and some are wrong. Measure both before training on them.

MetricAsks
CoverageDid a route run at all?
RecallWhen a specialist was needed, was it selected?
PrecisionWhen a route fired, was it correct?
LatencyDid the answer arrive while it was still useful?
Cost per successful taskIncluding retries, cache misses, and escalations

Rules for a fair comparison:

  • Compare against the human baseline, measured. If people route by hand today, the bar is their error rate, not perfection.
  • Agreement is not correctness. Precision against human labels measures agreement. Send disagreements to adjudication.
  • Evaluate whole sessions, held out from anything the router trained on.
  • Version the policy. Router version, threshold, candidate models, prompts, and fallback order are part of the configuration under test. An untracked policy change invalidates the comparison.
  • Replay history first. A harness that replays past traffic gives a baseline before any production change.

What survives a model swap: the harness, the labelled set, and the baseline. A tuned adapter is tied to a base model. Build the measurement first and keep it provider-agnostic.

A managed router from Fireworks, reviewed from public documentation.

AspectDocumented behaviour
GranularityChooses a model for each new user turn; that model handles the turn, including its tool calls
Model IDsfirerouter, a family alias such as firerouter/opus, or an exact pin such as firerouter/claude-opus-5-5
Custom routesUp to eight models, firerouter/{model1}/{model2}/...; the first is primary
Preferencesmax-intelligence, more-intelligence, balanced, more-savings, max-savings; support varies by harness
ReportingThe response model field names the model that served the request
BillingYou pay for the model that served each turn; closed models bill through your own provider credential; cached input bills at the serving model’s cached rate

What is not published, as reviewed: the features the router uses, its training data, and a measured comparison against the oracle. Read it as an instance of a session-level, cache-aware router, and apply section 10 before adopting it.

Notice the design choices against the sections above: per-turn granularity (section 5), cache awareness (section 6), and a user-set preference in place of a hidden threshold (section 2).

A category named on September 15, 2026, when TypeSafe AI announced Jev. The name borrows Kahneman’s fast, intuitive System 1. It is days old as of this review, so treat every number here as vendor-reported and unreplicated.

Definition. A decision model takes content plus a typed question and returns a typed answer, a probability for every option, and a confidence value. It does not write text. It produces all outputs in one forward pass instead of a token-by-token loop.

PrimitiveQuestion shapeReturns
ChoicePick one of a listed set, up to 255 optionsThe selection and a probability per option
ScorePlace content on an ordered scale of 2 to 10 levelsA probability-weighted position
NoulYes or noOne probability from 0 to 1

Why it belongs in a routing track. Every router in section 2 is a decision with a closed answer set. Three ways to make that decision, in order of how much machinery runs:

ApproachWhat runsWhat you get
LLM, free textFull generationA string to parse and validate
LLM, structured outputGeneration constrained to a schemaA valid shape; correctness not guaranteed
Decision modelOne pass, no generationA typed value and a probability to threshold on

Structured output guarantees the shape. It does not give a trained, calibrated probability. The claimed advantage of a decision model lives in the probabilities, not in the types.

Calibration is the product. Jev is described as trained with RLCD, a method that rewards honest probabilities: when the model says 0.8 it should be right about 80% of the time. Calibration is what makes confidence escalation from section 7 workable: act above a threshold, escalate below it. Measure it yourself with a reliability plot on your own labelled data before trusting a threshold.

What it cannot do, and what it still gets wrong.

  • It cannot invent a value outside your option list. It can pick the wrong one.
  • Documented weak spots for Jev 1.13 include arithmetic and counting, dates treated as text instead of ordered quantities, and literal reading of questions, where scoping words and double negatives change the answer.
  • It reads supplied text as data, and prompt injection is listed as a weakness. Section 7 still applies: on untrusted input, category rules outrank any probability.

The field, as listed by an independent directory.

ModelVendorAnnouncedAccess
Jev 1.13TypeSafe AISep 15, 2026Hosted API
Kev (0.8B, 4B, 9B)Jared PalmerSep 17, 2026Apache 2.0
LayaConvai InnovationsSep 18, 2026Apache 2.0
Tev1-4BTogether AISep 23, 2026Hosted API
CLM-8BContrastive-LMSep 23, 2026Apache 2.0
GLiNER2.5-Decide (340M)Fastino LabsSep 24, 2026Apache 2.0

Reported figures: Jev at $0.042 per million input tokens with output free, and 70 to 500 ms per decision. TypeSafe has not published Jev’s architecture, its training data, or the RLCD method. The open reproductions use their own training recipes, so they test the interface, not the method.

How to evaluate one. Same harness as section 10, plus calibration. Compare against two baselines you already own: a small LLM with structured output, and a fine-tuned encoder classifier. If a 300M-parameter classifier trained on your labels matches it, you did not need a new category.

13. What Goes in the Pool: Specialised Models

Section titled “13. What Goes in the Pool: Specialised Models”

A router is only as good as its candidates. The open-weight-only oracle in section 3 reached 90.3%, which says the cheap tier is no longer a fallback. The models filling that tier are open bases post-trained for one job. Figures below are vendor-reported, from Fireworks’ own posts.

ModelBaseHow it was builtReported result
Ember-1 (research preview, Sep 23, 2026)Kimi K3Over 50 training experiments and 200 evaluationsAbout 40% fewer tokens than the base at comparable quality; 75.2% against 66.4% on DeepSWE 1.1
Gen-1 Slides (with Genspark)MiniMax M3SFT on curated decks, then RL on a curriculum of growing context lengthLow-rated decks cut from 18% to 3.6%; about 90% lower cost per deck than Opus 5
dfs-large1 (depthfirst)GLM 5.2Post-trained with RLReported as a new Pareto frontier on a security benchmark
FARE-20Bgpt-oss-20BRejection-sampled SFT on a multi-domain setA multi-task evaluator: pairwise, step-level, reference-based and reference-free scoring
FireFunctionOpen baseAdditional training for tool useFunction calling

Four patterns repeat across that table:

  1. SFT first, RL second. Supervised data teaches the format; a reward teaches judgment. This is the ladder in section 9, applied to the worker instead of the router.
  2. Token efficiency is a training target. Ember-1 optimises how many tokens a task consumes, not only whether it succeeds. Cost per successful task falls even when accuracy is flat.
  3. The evaluator is a model too. FARE-20B exists because every reward and every eval needs a scorer. A tuned evaluator is infrastructure.
  4. Long-horizon RL is an engineering problem. The Gen-1 write-up names sequence-level rewards, alignment of tokenization between inference and training, and numerical stability on episodes over 100,000 tokens.

Benchmark scores overstate deployed value. The Specialized Intelligence Index cites a METR review of 296 AI-generated pull requests in which maintainer acceptance averaged 24.2 points below the automated benchmark score. The same caution as section 7 of the Grounded SQL Agent: a score is a capability claim, and a business outcome is a separate measurement.

Check how a specialised model is served. Several of these are on-demand only, which means dedicated capacity billed by GPU time. Section 8’s warning applies before any of them enters a per-token cost model.


SituationReach for
A deterministic feature already separates the trafficRules
Most requests are easy, a few need a strong modelPredictive router with a tuned threshold
A cheap check can verify the answerVerifier-gated cascade
Long agent sessions with large contextSession-level routing; price the cache
Untrusted content decides the routeCategory escalation, never confidence alone
No labelled routes yetPrompted classifier plus a replay harness
Labelled routes existSFT on them, LoRA first
A verifiable outcome and no gold answersRFT with your eval as the reward
A vendor shows an oracle chartAsk for the deployed router’s numbers on matched tasks
A closed-set decision at high volume, thresholded in codeA decision model or a tuned classifier; verify calibration
A narrow task with a verifiable reward and steady volumeA specialised post-trained model in the pool
  • Upper bounds are not results. An oracle, a best-case benchmark, and a vendor chart are all statements about opportunity.
  • Measure before you route. The baseline comes first, because improvement cannot be shown against a number nobody took.
  • Price the state you abandon. Cache, context, and consistency are costs of switching.
  • Put the control outside the model. Category rules and verifiers cannot be persuaded.
  • The labelled set is the asset. Models are replaced every few months; labels and harnesses are not.
  • Small, complementary, curated. Two or three well-chosen models beat a long list.
  • Match the machinery to the answer space. Closed set: decide. Open set: generate. Do not pay for generation to pick one of six labels.
ConceptConnected TrackApplication
Harness design, precision and recall, judgesLLM EvaluationMeasuring a routing policy
Prefill, KV cache, serving costLLM Systems & InferenceWhy a switch is expensive
Gateways, engines, provider selection, reproducibilityLLM Serving PlatformsThe layers a routed request passes through
SFT, DPO, LoRA, small modelsTraining & FrameworksTuning the router and the tiers
Retrieval over a large candidate setRetrieval & RAGRouting across hundreds of specialists is ranking
Deterministic gates, bounded retryGrounded SQL AgentVerifier-gated escalation in a real build
A checker proved by oracles and mutantsAgent Evaluation HarnessWhat a reward function has to survive before RFT trains on it
Least privilege, untrusted inputAuthorization & Access ControlWhy confidence is not a control
CompanyThe pattern they lean onInstance
FireworksSession-level, cache-aware routing across open and closed modelsFireRouter
OpenRouterGateway with provider selection, fallbacks, and optional model selectionAuto Router
LMSYS / AnyscalePreference-trained predictive routingRouteLLM
StanfordLearned cascades with a scoring functionFrugalGPT
TypeSafe AI / Together AITyped decisions with calibrated probabilitiesJev, Tev1
Fireworks with partnersOpen base, SFT then RL, against a domain rewardEmber-1, Gen-1 Slides
Any agent productCategory escalation and fail-open fan-outSpecialised subagents behind a classifier