Model Routing 101 - Fallback, Cost, and Latency
This is what Smart Inference does, and building it yourself is viable if your scale is small. Here is roughly how it works underneath, in case you want to build the simple version yourself before deciding whether you need the full thing.
The problem in one sentence
You have more than one place a request could go, more than one provider, more than one model, and you need to pick one per request based on some mix of cost, health, and latency, without hardcoding a single choice that goes stale the moment a provider has a bad day.
Step one: fallback chains
The simplest useful version is a fallback chain. Try provider A. If it fails or times out, try provider B. If that fails, try provider C.
This alone is worth building even if you do nothing else in this post. A single hardcoded provider with no fallback means one outage becomes your outage. A three-line fallback chain removes most of that risk for very little code.
Step two: health scoring instead of a fixed order
A static fallback chain has a problem: it always tries A first, even if A has been failing for the last hour. You end up eating A’s timeout on every single request before falling through to B, which is slow and wasteful.
Track a rolling health score per provider instead of a fixed order.
health_score can be as simple as error_rate * weight_a + p95_latency_ms * weight_b. The point is that the ordering itself becomes dynamic, a provider having a bad afternoon drops down the list without anyone touching a config file.
Step three: cost-weighted routing
Once you have health scoring, layer cost on top. Among the providers passing the health bar, pick the cheapest, not the first.
The trap here is estimating cost from a stale rate card. Providers change prices, sometimes hourly on the GPU side. If your estimated_cost function reads from a config file you update manually, you are back to the spreadsheet problem, a static number that goes wrong quietly. Pull current pricing from the provider’s API where possible, or measure actual billed cost from recent requests and use that as your estimate.
Step four: measuring what actually happened, not just what you predicted
This is the part that separates a routing prototype from something you can trust with real spend. Every response should get logged with what it actually cost, not what you expected it to cost.
Feed record_metric back into the health score and cost estimate for the next request. This closes the loop, the system learns from what actually happened instead of running on a fixed assumption forever.
When to build this yourself vs. not
If you have two providers and modest volume, the fallback chain from step one, maybe fifty lines of code, covers most of the real risk. Build it. You do not need a platform for this.
Where it gets genuinely annoying to maintain yourself: once you have more than a couple of providers, need per-GPU rather than per-VM cost tracking for self-hosted inference, or want the health scoring and cost data to survive a redeploy instead of resetting every time your process restarts. That is the point where Smart Inference starts saving real time, it is this same logic, already built, already tracking actual per-request cost instead of estimated, with the health and pricing data persisted instead of living in a process that restarts.
Either way, understand the mechanism before you reach for a tool that does it for you. It is not complicated. It is just tedious to get right and easy to let drift once it is working well enough to ignore.