Skip to content
EngineeringJuly 12, 2026

How automatic failover works in an LLM gateway

A single provider call has a single point of failure: one rate limit, one outage, one bad deploy on their end, and your request fails. Here's the actual mechanism AnyRouter uses to route around that — circuit breakers, escalating cooldowns, and a fallback chain — and why the extra hop barely registers next to model generation time.

One provider is a single point of failure

Call a provider's API directly and you've made a bet: that their endpoint stays up, that you're under their rate limit, and that nothing on their side degrades mid-request. Most of the time that bet pays off. When it doesn't, the failure is yours to handle — retry logic, a second SDK, a second key, code that only gets exercised during an incident.

A gateway exists to make that bet unnecessary. provider/model strings like z-ai/glm-5.2 or openai/gpt-4o-mini resolve to a list of upstream candidates, not one fixed destination, and the router picks among them per request based on live health.

The fallback chain

Every request goes through the same selection sequence before a single byte reaches an upstream:

  1. Select a primary candidate via the routing strategy
  2. Load every configured upstream for that model
  3. Apply request-level provider preferences, BYOK eligibility, and provider compliance policy
  4. Remove any backend currently in cooldown
  5. Return an ordered list — primary, then fallback 1, fallback 2, ...
  6. If every candidate is cooled down, return the full list anyway — degraded service beats none

The routing strategy that produces the primary candidate is configurable per request:

StrategyBehavior
latencyFastest upstream first
costCheapest upstream first
weighted_round_robinDistribute by configured weights
priorityTry upstreams in a fixed priority order
randomCrypto-secure random pick
ab_testDeterministic userId hash → variant

If the primary candidate's request fails, the executor doesn't surface the error — it moves to the next candidate in the list and tries again, transparently to the caller.

Circuit breakers stop hammering a broken upstream

Retrying a dead endpoint on every request wastes time and can make an incident worse. Each upstream backend has its own circuit breaker with fixed thresholds:

SettingValue
Failure threshold5 failures
Recovery timeout60s
Half-open max calls3
Success threshold to close2
Max retries2
Retry delay1s (exponential backoff)

Once a backend trips the breaker, it's skipped outright — no wasted round trip, no added latency for the caller — until the recovery timeout passes and the breaker allows a small number of half-open probe calls through.

Cooldowns escalate so healthy upstreams don't get starved

A circuit breaker is per-backend and short-lived. Underneath it, a KV-backed cooldown store tracks failure patterns over a rolling window — 3 failures inside a 60s observation window trips a cooldown — and each repeat offense doubles the penalty, starting at a 60s base and capping at 10 minutes:

Cooldown duration by escalation tier

Base 60s, doubles per tier, capped at 10 minutes

Tier 160s
Tier 2120s
Tier 3240s
Tier 4480s
Tier 5+ (cap)600s

Cooled-down backends are filtered out of the candidate list before routing even runs, so a flapping upstream doesn't get a steady trickle of production traffic while it recovers. In simplified form, the dispatch loop looks like this:

const candidates = await getFallbackChain(modelId, requestOptions)
  // primary + fallbacks, cooled-down backends already removed

for (const backend of candidates) {
  if (circuitBreaker.isOpen(backend)) continue

  try {
    const response = await dispatch(backend, requestBody)
    circuitBreaker.recordSuccess(backend)
    return response
  } catch (err) {
    circuitBreaker.recordFailure(backend)
    cooldownStore.recordFailure(backend) // may trigger escalating cooldown
    // fall through to the next candidate
  }
}
throw new AllUpstreamsFailedError(modelId)
Simplified — see packages/server-core/src/utils/upstream/executor.ts for the real implementation.

Why the extra hop is cheap

The honest answer here isn't a benchmark number — it's the shape of the two operations being compared. Selecting a candidate is a routing decision plus a cache read; generating a response is a model doing inference for as long as it takes to produce the output, one token at a time. Whatever the routing overhead is, it's happening once per request, while token generation happens for the entire duration of the response.

It also isn't paid for twice. The gateway doesn't buffer the upstream's response before returning it — chat completions, Anthropic messages, and Responses API streams are all piped through as server-sent events, translated dialect-to-dialect chunk by chunk, not accumulated into memory and replayed. Time-to-first-token is dominated by the upstream model, not by the routing layer in front of it.

AnyRouter live network stats showing aggregate tokens and requests processed through the gateway
This is the traffic the fallback chain has to hold up under — every request routed, none of it buffered client-side.

What this buys you

  • A model id keeps working through a provider outage — the request moves to the next configured upstream automatically
  • A struggling upstream stops receiving traffic within one failure window, not after a human notices
  • Recovery is gradual — half-open probe calls, not a full traffic dump the moment a backend looks healthy again
  • None of this requires retry logic in your client — it's the same request, same key, same response shape either way

Route through 170+ models across 28+ providers with automatic failover built in.

Open the dashboard

Route your first request in 2 minutes

Start free with your own keys, or top up and pay per token. Get $4/mo in credits and free models on Go — $2/mo, or free when you donate a provider key.

Start free