Skip to content

Smart Routing

How AnyRouter picks an upstream provider for each request — priority order, automatic failover, model aliases, and explicit override controls.

Pinning your app to a single provider means one outage or rate-limit spike takes you down with it. Every request to AnyRouter instead flows through a routing layer that picks the best upstream for the requested model — deterministic by default, respectful of health and cooldown state, and tunable per request when you need reproducibility or cost and latency controls.

Overview

You can select a model in three ways:

MethodExample
Direct modelanthropic/claude-sonnet-4.6
Preset reference@preset/code-reviewer
Preset with overrideopenai/gpt-5.4-mini@preset/code-reviewer

See Presets for full documentation on the preset syntax.

How it works

How a model maps to an upstream

AnyRouter maintains a list of upstream candidates for every model id. Each candidate can include routing metadata such as priority, price, latency, and throughput. On each request, AnyRouter builds the candidate set, applies your request-level provider preferences, skips cooled-down backends, and dispatches to the highest-priority remaining candidate. If that upstream fails with a retryable error, AnyRouter walks the remaining ordered candidates as fallbacks — unless you disabled fallbacks for that request.

flowchart TD
  A[Request: model id or preset] --> B[Build candidate set]
  B --> C[Apply provider preferences]
  C --> D[Skip cooled-down backends]
  D --> E[Dispatch to highest-priority candidate]
  E --> F{Success?}
  F -->|Yes| G[Return response]
  F -->|Retryable error| H{Fallbacks allowed?}
  H -->|Yes| I[Try next ordered candidate]
  I --> F
  H -->|No| J[Return error]

Request-level provider preferences

Steer routing directly in the request body with the provider object:

{
  "model": "openai/gpt-5.4-mini",
  "provider": {
    "only": ["OpenAI", "Groq"],
    "order": ["Groq", "OpenAI"],
    "sort": "latency",
    "allow_fallbacks": true,
    "max_price": {
      "prompt": "1.00",
      "completion": "2.00"
    },
    "preferred_max_latency": 300,
    "preferred_min_throughput": 40
  }
}

Supported controls:

  • only limits the candidate set to the named providers.
  • ignore removes named providers from the candidate set.
  • order pins a preferred backend order before the default route order.
  • sort reorders candidates by price, latency, throughput, or exacto.
  • require_params (for example ["tools"]) keeps only backends that support those parameters. Distinct from boolean require_parameters.
  • min_context is a token floor. anyrouter/auto drops members below it.
  • max_price.prompt and max_price.completion filter out backends above your budget.
  • preferred_max_latency drops candidates that exceed your latency target.
  • preferred_min_throughput drops candidates that fall below your throughput target.
  • allow_fallbacks: false disables retrying the next backend after the first candidate fails.

Provider names are normalized, so values like OpenAI, openai, Z.AI, and Moonshot AI map to the correct backend ids automatically.

Failover

A candidate is marked unhealthy for a short window when it returns a 5xx status, times out (timeout_ms exceeded), or exhausts its rate-limit budget. Unhealthy candidates are skipped until the cooldown expires. Requests that hit a fallback are still recorded with the final backend id, so you can see when failover happened.

Forcing a specific upstream

Set the X-AnyRouter-Provider header on any request to pin it to a single candidate:

curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer sk-ar-your-key" \
  -H "X-AnyRouter-Provider: openrouter" \
  -H "Content-Type: application/json" \
  -d '{"model": "openai/gpt-5.4-mini", "messages": [{"role": "user", "content": "hi"}]}'

Useful for:

  • Reproducibility — pin to a specific provider so your eval results don't silently drift when the router fails over.
  • BYOK tests — verify a specific BYOK credential is working end-to-end.
  • Outage debugging — route around a failing provider without waiting for the health cooldown.

If the forced provider is unhealthy, the request fails immediately. Header pinning is stricter than request-body preferences and bypasses normal fallback selection.

Model aliases

Some models have multiple slugs that resolve to the same upstream, for compatibility with other gateways. For example, anthropic/claude-sonnet-4.6 and claude-sonnet-4.6 both route to the same backend. Canonical slugs are provider/model; bare slugs are resolved by lookup.

AnyRouter also ships built-in virtual aliases under anyrouter/*. They appear in the models catalog and work on chat completions, messages, and responses:

AliasWhat it routes to
anyrouter/autoTop usage + top stable platform models (free-tier mixed with credit-billed), health-aware failover. Honors provider.sort: "exacto", min_context, and require_params: ["tools"]. anyrouter/auto[1m] / [500k] are min-context shortcuts, not separate listing ids.
anyrouter/latestThe latest stable production chat model that is actually serving traffic. Preview, experimental, and beta SKUs are skipped. Ranked from live request volume over the last 24 hours; when that usage rollup is unavailable, a documented fallback list of current platform-served flagships is used (google/gemini-3.5-flash, moonshotai/kimi-k2.6, z-ai/glm-5.2, x-ai/grok-4.5, then deepseek/deepseek-v4-pro and openai/gpt-4.1).
anyrouter/freeCurated $0 platform models (Go plan).
anyrouter/hermes / anyrouter/coworkPer-provider tool-calling chains, billed to credits.
anyrouter/byok / anyrouter/coding / anyrouter/agentChains built from your own provider keys.

anyrouter/latest is not "the newest YAML file in the catalog." A just-added preview SKU with no traffic never wins. Pin a concrete id when you need a specific generation; use anyrouter/latest when you want the current flagship that callers are already succeeding on.

Retired listing ids can keep working on chat completions after the SKU is removed from the catalog. The request is rewritten to a configured fallback (today stealth/ox-alpha and stealth/ox-alpha[1m] route to anyrouter/auto). The old model page stays up and says the model has been removed.

Tracing and request ids

Inference endpoints accept an optional request-body trace object:

{
  "trace": {
    "trace_id": "4f8c2b9670b44b49a2e71e7fd5c0b1d3",
    "trace_name": "checkout-flow",
    "span_name": "summarize-cart"
  }
}

When present, AnyRouter keeps that trace id attached to internal routing steps and forwards trace correlation metadata to AI Gateway-backed upstream calls. Every response also includes an X-Request-ID header so support can locate the exact request lifecycle.

Rate-limit headers

Every response includes rate-limit metadata for both the AnyRouter key and the upstream candidate that served the request:

X-RateLimit-Limit: 600
X-RateLimit-Remaining: 599
X-RateLimit-Reset: 1760000060
X-Upstream-RateLimit-Remaining: 998

Use these to back off gracefully — the router starts failing over if the selected candidate's upstream limit is close to exhaustion, but clients that respect the header get smoother behavior.

Related