AI Infrastructure · Operations
Multi-Model Routing in Practice
The pitch is obvious: send easy work to the cheap model. The part nobody mentions is that you now have a distributed system where quality varies by request.
Key takeaways
- Route on things you know before the call. Predicting difficulty per request is a research problem, not an infrastructure one.
- Cheap-model savings are real but bounded by retry amplification — a route that fails 30% of the time and escalates is not cheap.
- Tool-calling reliability differs more between models than raw reasoning quality. For agents, that difference dominates.
- Measure quality per route. An aggregate score hides a route that is quietly failing a fifth of its requests.
- Never route without a fallback that requires no model to decide. Routers fail; the fallback must not depend on the thing that failed.
The short answer
Route on task class, context length, latency budget, data residency and observed failure — not on a live quality prediction. Model routing genuinely reduces cost, often substantially, and it turns one dependency into several with different rate limits, tokenizers, tool-calling behaviours and failure shapes. Treat it as a capacity and cost decision with a quality floor, measure per-route outcome quality rather than aggregate, and always keep a deterministic fallback.
What routing is actually for
Three legitimate motivations, in descending order of how often they justify the complexity.
- 01CostMost requests in a production agent workload are not hard. Classification, extraction, routine summarisation and short tool-selection steps run acceptably on far cheaper models, and those steps are the bulk of the volume.
- 02LatencyInteractive paths need a fast response more than a marginally better one. Batch paths do not care. One model rarely sits at the right point on that curve for both.
- 03ConstraintData residency, contractual restrictions, vendor concentration risk and rate-limit ceilings all force multi-vendor whether or not you wanted the complexity.
The honest framing
Routing is capacity management with a quality floor. It is not a quality improvement. Any design that presents it as “always pick the best model for the job” is describing something that would require knowing the answer before asking.
Five signals worth routing on
Each of these is knowable before the call, cheap to evaluate, and stable enough to reason about in an incident.
Task class
Declared by the calling code, not inferred. Extraction, classification, drafting, tool selection and long-form reasoning have genuinely different requirements, and your application already knows which it is asking for.
Context length
A hard constraint, not a preference. Requests above a threshold have a shorter list of viable destinations, and this check is free.
Latency budget
Interactive versus batch, passed in by the caller. Batch work is where the deepest cost savings live because nobody is waiting.
Data residency and sensitivity
Region and classification of the data in the request. This is policy, not optimisation — it can only ever narrow the candidate set.
Observed failure, in retry
After a failure or a low-confidence result, escalate to a stronger model. This is the one quality-driven signal that works, because it is reactive rather than predictive.
Three that sound clever
A model that picks the model
Adds a full inference to every request, gets it wrong in ways nobody can debug, and creates a dependency whose failure takes down all routes at once. The router must be simpler than the thing it routes to.
Live quality prediction per prompt
Predicting whether a given request needs a stronger model requires understanding the request, which is the work itself. In practice these classifiers regress to prompt length.
Racing models and picking the best answer
Doubles or triples cost on every request to sometimes gain a little quality, and needs a judge — which is another model, with its own failure mode.
A rule that holds up
If your routing logic needs a model to run, you have added a single point of failure in front of your redundancy. Routing decisions should be resolvable from declared metadata and a lookup table.
What breaks
The interesting problems are not about quality. They are the ordinary consequences of turning one dependency into several.
| Surface | What differs | Consequence |
|---|---|---|
| Tokenizers | Different token counts for identical text | Cost and context estimates drift per route |
| Tool calling | Schema adherence and arity handling vary | An agent that works on one model loops on another |
| Rate limits | Separate ceilings, separate windows | Failover concentrates load and trips the fallback |
| Error shapes | Different codes, retry semantics, backoff advice | Retry logic written for one vendor misfires |
| Determinism | Temperature and seed handling differ | Reproducing an incident becomes harder |
| Safety filters | Different refusal thresholds | The same request succeeds on one route and refuses on another |
| Versioning | Models change under a stable alias | Quality shifts with no deploy on your side |
Scroll the table horizontally on narrow screens.
The tool-calling row is the one that bites agent systems hardest. Reasoning benchmarks do not predict whether a model reliably produces schema-valid tool calls with the right arity, and that behaviour is what an agent depends on for every step.
Measuring it honestly
- ›Aggregate quality score across all routes
- ›Average cost per request
- ›Model-level benchmark results
- ›Uptime of the router
- ›Outcome quality per route, tracked separately
- ›Cost per completed outcome, including retries
- ›Schema-valid tool-call rate per model
- ›Escalation rate from cheap to strong route
Escalation rate is the number that decides whether routing is working. If a third of cheap-route requests escalate, you are paying for both models plus latency, and the routing is a net loss. Track it per task class — the aggregate will hide the one class that never should have been routed down.
Cost per completed outcome, not per request, is the right denominator throughout. That framing and how to build it is covered in cost per outcome: a worked example.
- Keep a deterministic fallback route that requires no model to select.
- Pin model versions where the vendor allows it, and alert on alias changes.
- Replay a fixed evaluation set against every route weekly, not just at adoption.
- Record the chosen route in the execution record, or incidents become unreproducible.
The product
Route recommendation, health and cost in one control plane
Barzel Central Gateway compares routes by region, owner, protocol, cost, tool risk and policy constraints, and flags duplicate, missing or unhealthy paths. FinOps Atlas carries the cost attribution.
Key terms
- Model routing
- Directing a request to one of several models based on declared properties of the request.
- Escalation rate
- The share of requests sent to a cheaper route that subsequently require a stronger model.
- Retry amplification
- The multiplication of cost and latency caused by failures that trigger a second, more expensive attempt.
- Deterministic fallback
- A route selected without model inference, so it remains available when the routing layer fails.
- Task class
- A declared category of work — extraction, classification, drafting, tool selection, reasoning — used as a routing input.
Frequently asked questions
What should multi-model routing be based on?
Signals knowable before the call: declared task class, context length, latency budget, data residency and sensitivity, plus reactive escalation after an observed failure. Predictive per-request quality estimates do not work reliably.
Should a model choose which model to use?
No. It adds an inference to every request, fails in ways that are hard to debug, and puts a single point of failure in front of your redundancy. Routing should resolve from declared metadata and a lookup table.
Does model routing actually save money?
Usually yes, because most production requests are not hard. The saving is bounded by escalation: if a large share of cheap-route requests need a second attempt on a stronger model, you pay for both plus the latency.
What breaks when you route between models?
Tokenizer differences distort cost and context estimates, tool-calling reliability varies more than reasoning quality, rate limits and error shapes differ per vendor, and safety thresholds cause the same request to succeed on one route and refuse on another.
How should routing quality be measured?
Per route, never in aggregate: outcome quality by route, schema-valid tool-call rate per model, escalation rate per task class, and cost per completed outcome including retries.
Sources and further reading
Primary specifications and standards this article relies on. Where a claim is our own judgement rather than something a standard states, the article says so in the text.
- 01 · CNCF OpenTelemetry semantic conventions for GenAI ↗ Agreed attribute names for model, token and tool-call telemetry.
- 02 · CNCF OpenTelemetry ↗ Standard for the traces and spans an agent execution record should emit.
- 03 · FinOps Foundation FinOps Framework ↗ Capability model and phases referenced throughout.
- 04 · FinOps Foundation FOCUS — FinOps Open Cost and Usage Specification ↗ Open schema for normalising cost and usage data across providers.
- 05 · Google Google SRE Book — Service Level Objectives ↗ Where the SLI/SLO/error-budget vocabulary used here comes from.
- 06 · OWASP Foundation OWASP GenAI Security Project ↗ Ongoing community work on agentic and LLM security guidance.
Last reviewed 19 August 2026. External links open in a new tab; we do not control their content.
Reference documentation
Want the specification rather than the argument?
Route recommendation, comparison and routebooks — routing as a governed decision.
Free tier · 1,000 calls/mo · on the MCPize marketplace
Written by
Mark Alex
Founder of Real Biz Digital and architect of the Barzel ecosystem — five MCP servers published and callable in public. Software developer, technology entrepreneur and mechatronics engineer, working across AI agent governance, MCP security, AI infrastructure, FinOps and intelligent operations.