Model Fallbacks
Configure automatic failover between AI models when providers are down, rate-limited, or return retryable errors.
Keep a request alive across a full provider outage by listing backup models. The models parameter automatically tries other models if the primary model's providers are down, rate-limited, or otherwise return a retryable error.
This complements provider-level routing: provider routing fails over between upstreams for the same model, model fallbacks fail over between different models.
Before you start
- An AnyRouter API key (prefixed
sk-ar-). Create one in the dashboard. - Two or more model ids from the catalog to chain.
Model fallbacks are currently supported on /api/v1/chat/completions. The Responses and Messages endpoints don't read the models field yet — use /chat/completions for cross-model failover.
How it works
Provide an array of model IDs in priority order on the request body. AnyRouter:
- Resolves upstream candidates for the primary
modelfirst. - Resolves upstream candidates for each entry in
models, in order. - Concatenates the candidate lists and walks them with the same per-attempt timeout, retry, and cooldown logic used by single-model routing.
If model also appears in models, the duplicate is removed.
flowchart TD
A[Request: model + models] --> B[Try primary model's upstreams]
B --> C{Retryable failure?}
C -- No --> Z[Return response]
C -- Yes --> D[Try next model in models]
D --> E{Retryable failure?}
E -- No --> Z
E -- Yes --> F[Try next model...]
F --> G{List exhausted?}
G -- Yes --> H[Surface the error]
const response = await fetch("https://anyrouter.dev/api/v1/chat/completions", {
method: "POST",
headers: {
Authorization: "Bearer <ANYROUTER_API_KEY>",
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "anthropic/claude-sonnet-4.6",
models: ["openai/gpt-5.4-mini", "google/gemini-3.1-flash"],
messages: [
{ role: "user", content: "What is the meaning of life?" },
],
}),
})
const data = await response.json()
console.log(data.choices[0].message.content)
import json
import requests
response = requests.post(
"https://anyrouter.dev/api/v1/chat/completions",
headers={
"Authorization": "Bearer <ANYROUTER_API_KEY>",
"Content-Type": "application/json",
},
data=json.dumps({
"model": "anthropic/claude-sonnet-4.6",
"models": ["openai/gpt-5.4-mini", "google/gemini-3.1-flash"],
"messages": [
{"role": "user", "content": "What is the meaning of life?"}
],
}),
)
data = response.json()
print(data["choices"][0]["message"]["content"])
curl https://anyrouter.dev/api/v1/chat/completions \
-H "Authorization: Bearer <ANYROUTER_API_KEY>" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-4.6",
"models": ["openai/gpt-5.4-mini", "google/gemini-3.1-flash"],
"messages": [{"role": "user", "content": "What is the meaning of life?"}]
}'
The primary model is always tried first. Models listed in models are tried in array order. Entries in models that don't resolve to any configured upstream are skipped silently — you can pass an aspirational list and AnyRouter routes whatever it can.
Fallback Behavior
For each candidate, AnyRouter runs the upstream call with a 60-second per-attempt timeout. The next candidate is tried when the previous attempt returns a retryable HTTP status, times out, or aborts. Currently retryable: 402, 403, 404, 408, 425, 429, 500, 502, 503, 504, and network/abort errors. All other status codes (notably 400-level validation errors such as context-length overflow) surface immediately without falling back.
Within each model in the chain, AnyRouter still applies its normal provider routing — every eligible upstream for that model is attempted before moving on to the next model in models. Set provider.allow_fallbacks: false (combined with provider.order) to constrain the per-model candidate list to a single provider before falling over to the next model.
Fallback only happens on the initial response. Once an upstream returns 200 and AnyRouter starts piping SSE bytes back to the client, the stream cannot be rewound — mid-stream failures surface as-is.
Pricing
Requests are priced using the model that actually served the response. The response body's model field is rewritten to that catalog model id so your client can see, and your billing/usage row records, which model in the chain was billed.
{
"id": "chatcmpl-...",
"model": "openai/gpt-5.4-mini",
"choices": [...]
}
In this example the primary anthropic/claude-sonnet-4.6 request failed and AnyRouter billed openai/gpt-5.4-mini — the model that actually answered. The same value appears on the model field of the metadata frame emitted at the end of a streamed response.
When the upstream provider returns its own cost signal (via a Cloudflare AI Gateway response header or a usage.cost field on the response body), AnyRouter captures it and surfaces it alongside our locally-computed price. You'll see:
usage.cost— what AnyRouter charged you, computed from the served model's catalog pricing.usage.cost_details.upstream_inference_cost— what the upstream provider reported (when available); otherwise the same value asusage.cost.anyrouter_metadata.cost.upstreamAmount— the upstream-reported cost on the metadata frame for streamed responses (when available).
Using With the OpenAI SDK
The OpenAI SDK doesn't expose a top-level models field, so pass it via extra_body (Python) or as a raw additional property (TypeScript). The example below tries openai/gpt-5.4 first, then walks the models list.
from openai import OpenAI
client = OpenAI(
base_url="https://anyrouter.dev/api/v1",
api_key="<ANYROUTER_API_KEY>",
)
completion = client.chat.completions.create(
model="openai/gpt-5.4",
extra_body={
"models": ["anthropic/claude-sonnet-4.6", "google/gemini-3.1-flash"],
},
messages=[
{"role": "user", "content": "What is the meaning of life?"}
],
)
print(completion.choices[0].message.content)
import OpenAI from "openai"
const client = new OpenAI({
baseURL: "https://anyrouter.dev/api/v1",
apiKey: "<ANYROUTER_API_KEY>",
})
const completion = await client.chat.completions.create({
model: "openai/gpt-5.4",
// @ts-expect-error - non-standard OpenAI parameter, AnyRouter extension
models: ["anthropic/claude-sonnet-4.6", "google/gemini-3.1-flash"],
messages: [
{ role: "user", content: "What is the meaning of life?" },
],
})
console.log(completion.choices[0].message.content)
Choosing a Fallback Chain
A few patterns we see in production:
- Same-family redundancy — pair Claude with another Claude tier (
anthropic/claude-sonnet-4.6→anthropic/claude-haiku-4.5) when behavior parity matters more than provider diversity. - Cross-provider redundancy — pair Claude with GPT and Gemini to survive a full provider outage.
- Quality → speed degradation — start with a stronger model and fall back to a faster, cheaper one when load spikes (e.g.
openai/gpt-5.4→openai/gpt-5.4-mini). - Open-weight backstop — fall back to an open-weight model hosted on multiple providers (e.g.
meta-llama/llama-3.3-70b-instruct) to survive proprietary-provider outages.
Combining With Provider Sorting
When you want AnyRouter to pick the single best endpoint across all models (instead of strictly trying the primary first), combine models with provider.sort and partition: "none". See Advanced Sorting with Partition.
curl https://anyrouter.dev/api/v1/chat/completions \
-H "Authorization: Bearer <ANYROUTER_API_KEY>" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-4.6",
"models": ["openai/gpt-5.4-mini", "google/gemini-3.1-flash"],
"messages": [{"role": "user", "content": "Hello"}],
"provider": {
"sort": { "by": "throughput", "partition": "none" }
}
}'
With partition: "none", AnyRouter sorts every endpoint across all three models by throughput and tries the fastest one first — useful when you have multiple acceptable models and just want whichever performs best right now.
Verify
Send a request whose primary model is unlikely to be your cheapest route and inspect the returned model field — it tells you which model in the chain actually answered:
curl https://anyrouter.dev/api/v1/chat/completions \
-H "Authorization: Bearer <ANYROUTER_API_KEY>" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-4.6",
"models": ["openai/gpt-5.4-mini", "google/gemini-3.1-flash"],
"messages": [{"role": "user", "content": "Say hello."}]
}'
Troubleshooting
My fallback model was never tried
Fallback only triggers on retryable statuses — 402, 403, 404, 408, 425, 429, 500, 502, 503, 504, and network/abort errors. A 400-level validation error (like context-length overflow) surfaces immediately without falling back. Fix the request rather than expecting a fallback.
One of my models in the list is silently ignored
Entries in models that don't resolve to any configured upstream are skipped silently — you can pass an aspirational list and AnyRouter routes whatever it can. Confirm each id exists in Models.
My stream failed halfway and didn't fall back
Fallback happens only on the initial response. Once bytes are streaming, the stream can't be rewound. Add client-side retry logic for mid-stream failures.