One provider is a single point of failure
Call a provider's API directly and you've made a bet: that their endpoint stays up, that you're under their rate limit, and that nothing on their side degrades mid-request. Most of the time that bet pays off. When it doesn't, the failure is yours to handle — retry logic, a second SDK, a second key, code that only gets exercised during an incident.
A gateway exists to make that bet unnecessary. provider/model strings like z-ai/glm-5.2 or openai/gpt-4o-mini resolve to a list of upstream candidates, not one fixed destination, and the router picks among them per request based on live health.
The fallback chain
Every request goes through the same selection sequence before a single byte reaches an upstream:
- Select a primary candidate via the routing strategy
- Load every configured upstream for that model
- Apply request-level provider preferences, BYOK eligibility, and provider compliance policy
- Remove any backend currently in cooldown
- Return an ordered list — primary, then fallback 1, fallback 2, ...
- If every candidate is cooled down, return the full list anyway — degraded service beats none
The routing strategy that produces the primary candidate is configurable per request:
| Strategy | Behavior |
|---|---|
| latency | Fastest upstream first |
| cost | Cheapest upstream first |
| weighted_round_robin | Distribute by configured weights |
| priority | Try upstreams in a fixed priority order |
| random | Crypto-secure random pick |
| ab_test | Deterministic userId hash → variant |
If the primary candidate's request fails, the executor doesn't surface the error — it moves to the next candidate in the list and tries again, transparently to the caller.
Circuit breakers stop hammering a broken upstream
Retrying a dead endpoint on every request wastes time and can make an incident worse. Each upstream backend has its own circuit breaker with fixed thresholds:
| Setting | Value |
|---|---|
| Failure threshold | 5 failures |
| Recovery timeout | 60s |
| Half-open max calls | 3 |
| Success threshold to close | 2 |
| Max retries | 2 |
| Retry delay | 1s (exponential backoff) |
Once a backend trips the breaker, it's skipped outright — no wasted round trip, no added latency for the caller — until the recovery timeout passes and the breaker allows a small number of half-open probe calls through.
Cooldowns escalate so healthy upstreams don't get starved
A circuit breaker is per-backend and short-lived. Underneath it, a KV-backed cooldown store tracks failure patterns over a rolling window — 3 failures inside a 60s observation window trips a cooldown — and each repeat offense doubles the penalty, starting at a 60s base and capping at 10 minutes:
Cooldown duration by escalation tier
Base 60s, doubles per tier, capped at 10 minutes
Cooled-down backends are filtered out of the candidate list before routing even runs, so a flapping upstream doesn't get a steady trickle of production traffic while it recovers. In simplified form, the dispatch loop looks like this:
const candidates = await getFallbackChain(modelId, requestOptions)
// primary + fallbacks, cooled-down backends already removed
for (const backend of candidates) {
if (circuitBreaker.isOpen(backend)) continue
try {
const response = await dispatch(backend, requestBody)
circuitBreaker.recordSuccess(backend)
return response
} catch (err) {
circuitBreaker.recordFailure(backend)
cooldownStore.recordFailure(backend) // may trigger escalating cooldown
// fall through to the next candidate
}
}
throw new AllUpstreamsFailedError(modelId)Why the extra hop is cheap
The honest answer here isn't a benchmark number — it's the shape of the two operations being compared. Selecting a candidate is a routing decision plus a cache read; generating a response is a model doing inference for as long as it takes to produce the output, one token at a time. Whatever the routing overhead is, it's happening once per request, while token generation happens for the entire duration of the response.
It also isn't paid for twice. The gateway doesn't buffer the upstream's response before returning it — chat completions, Anthropic messages, and Responses API streams are all piped through as server-sent events, translated dialect-to-dialect chunk by chunk, not accumulated into memory and replayed. Time-to-first-token is dominated by the upstream model, not by the routing layer in front of it.

What this buys you
- A model id keeps working through a provider outage — the request moves to the next configured upstream automatically
- A struggling upstream stops receiving traffic within one failure window, not after a human notices
- Recovery is gradual — half-open probe calls, not a full traffic dump the moment a backend looks healthy again
- None of this requires retry logic in your client — it's the same request, same key, same response shape either way
Route through 170+ models across 28+ providers with automatic failover built in.
Open the dashboardRoute your first request in 2 minutes
Start free with your own keys, or top up and pay per token. Get $4/mo in credits and free models on Go — $2/mo, or free when you donate a provider key.
Start free