Smart Routing
How AnyRouter picks an upstream provider for each request — priority order, automatic failover, model aliases, and explicit override controls.
Pinning your app to a single provider means one outage or rate-limit spike takes you down with it. Every request to AnyRouter instead flows through a routing layer that picks the best upstream for the requested model — deterministic by default, respectful of health and cooldown state, and tunable per request when you need reproducibility or cost and latency controls.
Overview
You can select a model in three ways:
| Method | Example |
|---|---|
| Direct model | anthropic/claude-sonnet-4.6 |
| Preset reference | @preset/code-reviewer |
| Preset with override | openai/gpt-5.4-mini@preset/code-reviewer |
See Presets for full documentation on the preset syntax.
How it works
How a model maps to an upstream
AnyRouter maintains a list of upstream candidates for every model id. Each candidate can include routing metadata such as priority, price, latency, and throughput. On each request, AnyRouter builds the candidate set, applies your request-level provider preferences, skips cooled-down backends, and dispatches to the highest-priority remaining candidate. If that upstream fails with a retryable error, AnyRouter walks the remaining ordered candidates as fallbacks — unless you disabled fallbacks for that request.
flowchart TD
A[Request: model id or preset] --> B[Build candidate set]
B --> C[Apply provider preferences]
C --> D[Skip cooled-down backends]
D --> E[Dispatch to highest-priority candidate]
E --> F{Success?}
F -->|Yes| G[Return response]
F -->|Retryable error| H{Fallbacks allowed?}
H -->|Yes| I[Try next ordered candidate]
I --> F
H -->|No| J[Return error]
Request-level provider preferences
Steer routing directly in the request body with the provider object:
{
"model": "openai/gpt-5.4-mini",
"provider": {
"only": ["OpenAI", "Groq"],
"order": ["Groq", "OpenAI"],
"sort": "latency",
"allow_fallbacks": true,
"max_price": {
"prompt": "1.00",
"completion": "2.00"
},
"preferred_max_latency": 300,
"preferred_min_throughput": 40
}
}
Supported controls:
onlylimits the candidate set to the named providers.ignoreremoves named providers from the candidate set.orderpins a preferred backend order before the default route order.sortreorders candidates byprice,latency,throughput, orexacto.require_params(for example["tools"]) keeps only backends that support those parameters. Distinct from booleanrequire_parameters.min_contextis a token floor.anyrouter/autodrops members below it.max_price.promptandmax_price.completionfilter out backends above your budget.preferred_max_latencydrops candidates that exceed your latency target.preferred_min_throughputdrops candidates that fall below your throughput target.allow_fallbacks: falsedisables retrying the next backend after the first candidate fails.
Provider names are normalized, so values like OpenAI, openai, Z.AI, and Moonshot AI map to the correct backend ids automatically.
Failover
A candidate is marked unhealthy for a short window when it returns a 5xx status, times out (timeout_ms exceeded), or exhausts its rate-limit budget. Unhealthy candidates are skipped until the cooldown expires. Requests that hit a fallback are still recorded with the final backend id, so you can see when failover happened.
Forcing a specific upstream
Set the X-AnyRouter-Provider header on any request to pin it to a single candidate:
curl https://anyrouter.dev/api/v1/chat/completions \
-H "Authorization: Bearer sk-ar-your-key" \
-H "X-AnyRouter-Provider: openrouter" \
-H "Content-Type: application/json" \
-d '{"model": "openai/gpt-5.4-mini", "messages": [{"role": "user", "content": "hi"}]}'
Useful for:
- Reproducibility — pin to a specific provider so your eval results don't silently drift when the router fails over.
- BYOK tests — verify a specific BYOK credential is working end-to-end.
- Outage debugging — route around a failing provider without waiting for the health cooldown.
If the forced provider is unhealthy, the request fails immediately. Header pinning is stricter than request-body preferences and bypasses normal fallback selection.
Model aliases
Some models have multiple slugs that resolve to the same upstream, for compatibility with other gateways. For example, anthropic/claude-sonnet-4.6 and claude-sonnet-4.6 both route to the same backend. Canonical slugs are provider/model; bare slugs are resolved by lookup.
AnyRouter also ships built-in virtual aliases under anyrouter/*. They appear in the models catalog and work on chat completions, messages, and responses:
| Alias | What it routes to |
|---|---|
anyrouter/auto | Top usage + top stable platform models (free-tier mixed with credit-billed), health-aware failover. Honors provider.sort: "exacto", min_context, and require_params: ["tools"]. anyrouter/auto[1m] / [500k] are min-context shortcuts, not separate listing ids. |
anyrouter/latest | The latest stable production chat model that is actually serving traffic. Preview, experimental, and beta SKUs are skipped. Ranked from live request volume over the last 24 hours; when that usage rollup is unavailable, a documented fallback list of current platform-served flagships is used (google/gemini-3.5-flash, moonshotai/kimi-k2.6, z-ai/glm-5.2, x-ai/grok-4.5, then deepseek/deepseek-v4-pro and openai/gpt-4.1). |
anyrouter/free | Curated $0 platform models (Go plan). |
anyrouter/hermes / anyrouter/cowork | Per-provider tool-calling chains, billed to credits. |
anyrouter/byok / anyrouter/coding / anyrouter/agent | Chains built from your own provider keys. |
anyrouter/latest is not "the newest YAML file in the catalog." A just-added preview SKU with no traffic never wins. Pin a concrete id when you need a specific generation; use anyrouter/latest when you want the current flagship that callers are already succeeding on.
Retired listing ids can keep working on chat completions after the SKU is removed from the catalog. The request is rewritten to a configured fallback (today stealth/ox-alpha and stealth/ox-alpha[1m] route to anyrouter/auto). The old model page stays up and says the model has been removed.
Tracing and request ids
Inference endpoints accept an optional request-body trace object:
{
"trace": {
"trace_id": "4f8c2b9670b44b49a2e71e7fd5c0b1d3",
"trace_name": "checkout-flow",
"span_name": "summarize-cart"
}
}
When present, AnyRouter keeps that trace id attached to internal routing steps and forwards trace correlation metadata to AI Gateway-backed upstream calls. Every response also includes an X-Request-ID header so support can locate the exact request lifecycle.
Rate-limit headers
Every response includes rate-limit metadata for both the AnyRouter key and the upstream candidate that served the request:
X-RateLimit-Limit: 600
X-RateLimit-Remaining: 599
X-RateLimit-Reset: 1760000060
X-Upstream-RateLimit-Remaining: 998
Use these to back off gracefully — the router starts failing over if the selected candidate's upstream limit is close to exhaustion, but clients that respect the header get smoother behavior.