Cut LLM Costs with Automatic Provider Fallback
Route one model id across multiple upstreams so AnyRouter picks the cheapest healthy provider and falls back automatically when one is down.
Many models are served by more than one provider, often at different prices and with different rate limits. Let AnyRouter pick the best upstream for each request and fall back automatically when one is unavailable — with a single, unchanged model id in your code.
The idea
For a model like openai/gpt-oss-120b that several providers host, AnyRouter keeps a live view of price, latency, and health for each upstream. You express a preference — cheapest first, fastest first, or an explicit order — and AnyRouter honors it while skipping any provider that is currently failing. The routing happens on AnyRouter's side, so your request body barely changes.
flowchart LR
app["Your app"] --> ar["AnyRouter"]
ar -->|"1st choice (cheapest healthy)"| p1["Provider A"]
ar -.->|"fallback if A is down/limited"| p2["Provider B"]
ar -.->|"fallback"| p3["Provider C"]
How it works
You add a provider block to any Chat Completions request. sort controls preference, order pins an explicit sequence, max_price caps what you'll pay, and allow_fallbacks decides whether AnyRouter may move on when an upstream errors or is rate-limited. Prices throughout are in US dollars per million tokens.
Implementation
Sort by price
sort: "price" tries the cheapest healthy upstream first; allow_fallbacks: true lets it move on if that upstream errors or is rate-limited.
curl https://anyrouter.dev/api/v1/chat/completions \
-H "Authorization: Bearer sk-ar-v1-your-key" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-oss-120b",
"messages": [{"role": "user", "content": "Summarize this changelog in 3 bullets."}],
"provider": {
"sort": "price",
"allow_fallbacks": true
}
}'
Pin an explicit order (optional)
To control the order yourself — for example, prefer a provider where you have a volume discount — use order:
{
"model": "openai/gpt-oss-120b",
"messages": [{ "role": "user", "content": "Hello" }],
"provider": {
"order": ["deepinfra", "cerebras"],
"allow_fallbacks": true
}
}
AnyRouter tries deepinfra first, then cerebras, then any other healthy upstream (because allow_fallbacks is on). Set allow_fallbacks: false to fail fast and only use the providers you listed.
Cap the price you'll pay (optional)
To guarantee a request never exceeds a per-token budget, set max_price. Any upstream above the cap is skipped:
{
"model": "openai/gpt-oss-120b",
"messages": [{ "role": "user", "content": "Hello" }],
"provider": {
"sort": "price",
"max_price": { "prompt": 0.5, "completion": 1.5 },
"allow_fallbacks": true
}
}
If no upstream is under the cap, the request returns an error rather than silently picking a pricier provider.
Don't want to send a provider block on every call? Save the same preferences once as a preset and use the preset slug as the model. Every request through that slug inherits the routing rules.
Related
- Smart Routing — how AnyRouter chooses an upstream
- Model Fallbacks — cascade across multiple models
- Provider Routing — the full
providerblock reference - Chat Completions API — request and response shape