Skip to content

Provider Routing

Route requests across upstream providers. Customize ordering, sorting, fallbacks, performance thresholds, data policies, and pricing budgets per request.

Control which upstream provider serves each request — pin an order, cap price, prefer latency or throughput, or restrict to providers that meet your data policy. AnyRouter routes each request to the best available upstream for the model you asked for; the provider object lets you steer that per request.

By default, requests are load balanced across the top-priority providers to maximize uptime. You customize routing with the provider object in the request body for Chat Completions; the same object is accepted on Responses and Messages endpoints.

Before you start

  • An AnyRouter API key (prefixed sk-ar-). Create one in the dashboard.
  • A model id from the catalog to route.

Steer a request

Add a provider object to the body

Every routing control lives under a top-level provider key alongside model and messages.

Pick a strategy

Set sort to "price", "throughput", or "latency", or set order to pin a provider preference list. Add max_price to enforce a budget ceiling.

Decide on fallbacks

Leave allow_fallbacks at its default true for resilience, or set it to false when you need a request served only by your chosen provider.

Model-id shortcuts save you a provider object: append :nitro to sort by throughput, :floor to sort by price, :fast to sort by latency, or :exacto to prefer higher-precision weights (and quality-first members on anyrouter/auto). On anyrouter/auto, append [1m] or [500k] to keep only members with at least that context window (provider.min_context).

Field reference

The provider object can contain the following fields:

FieldTypeDefaultDescription
orderstring[]–List of provider slugs to try in order (e.g. ["anthropic", "openai"]). Learn more
allow_fallbacksbooleantrueWhether to allow backup providers when the primary is unavailable. Learn more
require_parametersbooleanfalseOnly use providers that support all parameters in your request. Learn more
require_paramsstring[]–Hard parameter requirements (for example ["tools"]). Distinct from require_parameters. Learn more
min_contextnumber–Minimum context window in tokens. anyrouter/auto keeps only members at or above this floor. Learn more
data_collection"allow" | "deny""allow"Control whether to use providers that may store data. Learn more
zdrboolean–Restrict routing to only Zero Data Retention endpoints. Learn more
enforce_distillable_textboolean–Restrict routing to only models that allow text distillation. Learn more
onlystring[]–List of provider slugs to allow for this request. Learn more
ignorestring[]–List of provider slugs to skip for this request. Learn more
quantizationsstring[]–List of quantization levels to filter by (e.g. ["int4", "int8"]). Learn more
sortstring | object–Sort providers by price, throughput, latency, or exacto. Can be a string (e.g. "price") or an object with by and partition fields. Learn more
preferred_min_throughputnumber | object–Preferred minimum throughput (tokens/sec). Number, or object with percentile cutoffs (p50/p75/p90/p99). Learn more
preferred_max_latencynumber | object–Preferred maximum latency (seconds). Number, or object with percentile cutoffs (p50/p75/p90/p99). Learn more
max_priceobject–The maximum pricing you want to pay for this request. Learn more

Support status. Every field documented on this page is honored end-to-end: order, only, ignore, allow_fallbacks, sort (string form including "exacto", partition: "model", and partition: "none" across models), max_price, preferred_max_latency, preferred_min_throughput, require_parameters, require_params, min_context, data_collection, zdr, quantizations, enforce_distillable_text, and the :nitro / :floor / :fast / :exacto model-id shortcuts. Filters that depend on provider metadata (zdr, quantizations, data_collection, enforce_distillable_text) operate on the per-upstream features block and the model's data_policy declarations in our catalog. As that metadata fills in across more providers, the result set for these filters will grow without any client change.

Price-Based Load Balancing (Default Strategy)

For each model in your request, AnyRouter's default behavior is to load balance requests across providers, prioritizing price.

If you are more sensitive to throughput than price, use the sort field to explicitly prioritize throughput.

When you send a request with tools or tool_choice, AnyRouter only routes to providers that support tool use. Similarly, if you set max_tokens, AnyRouter only routes to providers that support a response of that length.

The default load balancing strategy:

  1. Prioritize providers that have not seen significant outages in the last 30 seconds.
  2. For the stable providers, look at the lowest-cost candidates and pick one weighted by the inverse square of the price (example below).
  3. Use the remaining providers as fallbacks.
flowchart TD
  A[Request] --> B{Stable in<br/>last 30s?}
  B -- No --> F[Move to fallback pool]
  B -- Yes --> C[Weight by inverse-square of price]
  C --> D[Pick a primary provider]
  D --> E{Succeeds?}
  E -- Yes --> G[Return response]
  E -- No --> F
  F --> H[Try next candidate]
  H --> E

Example. If Provider A costs $1/M tokens, Provider B costs $2, and Provider C costs $3, and Provider B recently saw a few outages:

  • Your request is 9× more likely to land on Provider A than Provider C, because $(1 / 3^2 = 1/9)$.
  • If Provider A fails, Provider C is tried next.
  • If Provider C also fails, Provider B is tried last.

If you have sort or order set, load balancing is disabled.

Provider Sorting

To explicitly prioritize a particular provider attribute, include sort in the provider preferences. Load balancing is disabled, and the router tries providers in order.

The sort options are:

  • "price": prioritize lowest price
  • "throughput": prioritize highest throughput
  • "latency": prioritize lowest latency
  • "exacto": prefer higher-precision weights on a single model, and quality-first members on anyrouter/auto
fetch("https://anyrouter.dev/api/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: "Bearer <ANYROUTER_API_KEY>",
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model: "meta-llama/llama-3.3-70b-instruct",
    messages: [{ role: "user", content: "Hello" }],
    provider: { sort: "throughput" },
  }),
})
import requests

requests.post(
    "https://anyrouter.dev/api/v1/chat/completions",
    headers={
        "Authorization": "Bearer <ANYROUTER_API_KEY>",
        "Content-Type": "application/json",
    },
    json={
        "model": "meta-llama/llama-3.3-70b-instruct",
        "messages": [{"role": "user", "content": "Hello"}],
        "provider": {"sort": "throughput"},
    },
)

To always prioritize low prices, set sort to "price". To always prioritize low latency, set sort to "latency".

Nitro Shortcut

Append :nitro to any model slug as a shortcut to sort by throughput. This is exactly equivalent to setting provider.sort to "throughput".

curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer <ANYROUTER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/llama-3.3-70b-instruct:nitro",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

Floor Price Shortcut

Append :floor to any model slug as a shortcut to sort by price. This is exactly equivalent to setting provider.sort to "price".

curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer <ANYROUTER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/llama-3.3-70b-instruct:floor",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

Exacto Shortcut

Append :exacto to any model slug as a shortcut to set provider.sort to "exacto". On a concrete model this prefers higher-precision weights. On anyrouter/auto it picks members by quality instead of live usage. Exacto is a routing preference, not a new model id.

curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer <ANYROUTER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anyrouter/auto:exacto",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

Context-window shortcut

Append [1m] or [500k] (or any [<n>k] / [<n>m]) to anyrouter/auto to set provider.min_context. [1m] keeps members with at least 1M tokens; [500k] keeps members with at least 500k. Combine with :exacto (anyrouter/auto[1m]:exacto).

anyr claude --yolo --model anyrouter/auto[1m]
anyr claude --model anyrouter/auto[500k]

Advanced Sorting with Partition

When using model fallbacks, sort can be specified as an object with additional options to control how endpoints are sorted across multiple models.

FieldTypeDefaultDescription
sort.bystring–The sorting strategy: "price", "throughput", or "latency".
sort.partitionstring"model"How to group endpoints for sorting: "model" (default) or "none".

By default, when you specify multiple models (fallbacks), AnyRouter groups endpoints by model before sorting. This means the primary model's endpoints are always tried first, regardless of their performance characteristics. Setting partition to "none" removes this grouping, allowing endpoints to be sorted globally across all models.

preferred_max_latency and preferred_min_throughput do not guarantee you will get a provider or model with this performance level. Providers that hit your thresholds are preferred; those that don't are deprioritized as fallbacks. This is different from max_price, which is a hard filter — a request can fail if no provider meets it.

Use Case 1: Route to the Highest Throughput or Lowest Latency Model

When you have multiple acceptable models and want whichever has the best performance right now, use partition: "none" with throughput or latency sorting:

curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer <ANYROUTER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "models": [
      "anthropic/claude-sonnet-4.6",
      "openai/gpt-5.4-mini",
      "google/gemini-3.1-flash"
    ],
    "messages": [{"role": "user", "content": "Hello"}],
    "provider": {
      "sort": { "by": "throughput", "partition": "none" }
    }
  }'

AnyRouter routes to whichever endpoint across all three models currently has the highest throughput, rather than always trying Claude first.

Performance Thresholds

You can set minimum throughput or maximum latency thresholds to filter endpoints. Endpoints that don't meet these thresholds are deprioritized (moved to the end of the list) rather than excluded entirely.

FieldTypeDefaultDescription
preferred_min_throughputnumber | object–Preferred minimum throughput in tokens/sec. Number (applies to p50) or object with percentile cutoffs.
preferred_max_latencynumber | object–Preferred maximum latency in seconds. Number (applies to p50) or object with percentile cutoffs.

How Percentiles Work

AnyRouter tracks latency and throughput for each model and provider using percentile statistics calculated over a rolling 5-minute window:

  • p50 (median): 50% of requests perform better than this value
  • p75: 75% of requests perform better than this value
  • p90: 90% of requests perform better than this value
  • p99: 99% of requests perform better than this value

Higher percentiles (p90, p99) give more confidence about worst-case performance; lower percentiles (p50) reflect typical performance. If a provider has a p90 latency of 2 seconds, 90% of its requests complete in under 2 seconds.

When you specify multiple percentile cutoffs, all of them must be met for a provider to land in the preferred group.

When to Use Percentile Preferences

  • Real-time applications: use p90 or p99 latency thresholds for consistent user-facing response times.
  • Batch processing: use p50 throughput thresholds when average performance matters more than worst-case.
  • SLA compliance: use multiple percentile cutoffs to enforce performance tiers.
  • Cost optimization: combine with sort: "price" to get the cheapest provider that still meets your performance bar.

Use Case 2: Find the Cheapest Model Meeting Performance Requirements

curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer <ANYROUTER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "models": [
      "anthropic/claude-sonnet-4.6",
      "openai/gpt-5.4-mini",
      "google/gemini-3.1-flash"
    ],
    "messages": [{"role": "user", "content": "Hello"}],
    "provider": {
      "sort": { "by": "price", "partition": "none" },
      "preferred_min_throughput": { "p90": 50 }
    }
  }'

AnyRouter picks the cheapest model+provider across the three options that has at least 50 tokens/second throughput at the p90 level. Endpoints below this threshold are still available as fallbacks if all preferred options fail.

You can also use preferred_max_latency:

curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer <ANYROUTER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "models": ["anthropic/claude-sonnet-4.6", "openai/gpt-5.4-mini"],
    "messages": [{"role": "user", "content": "Hello"}],
    "provider": {
      "sort": { "by": "price", "partition": "none" },
      "preferred_max_latency": { "p90": 3 }
    }
  }'

Example: Multiple Percentile Cutoffs

curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer <ANYROUTER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek/deepseek-v3.2",
    "messages": [{"role": "user", "content": "Hello"}],
    "provider": {
      "preferred_max_latency":   { "p50": 1, "p90": 3, "p99": 5 },
      "preferred_min_throughput": { "p50": 100, "p90": 50 }
    }
  }'

Use Case 3: Maximize BYOK Usage Across Models

If you use Bring Your Own Key and want to maximize use of your own provider keys, partition: "none" helps. When your primary model has no BYOK provider configured, AnyRouter can route to a fallback model that does.

curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer <ANYROUTER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "models": [
      "anthropic/claude-sonnet-4.6",
      "openai/gpt-5.4-mini",
      "google/gemini-3.1-flash"
    ],
    "messages": [{"role": "user", "content": "Hello"}],
    "provider": { "sort": { "by": "price", "partition": "none" } }
  }'

If you have a BYOK key configured for OpenAI but not for Anthropic, AnyRouter can route to GPT using your own key even though Claude is listed first. Without partition: "none", the router always tries Claude's endpoints first before falling back.

BYOK endpoints are automatically prioritized when you have API keys configured for that provider. The partition: "none" setting allows this prioritization to work across model boundaries.

Ordering Specific Providers

Set the providers AnyRouter will prioritize using the order field.

FieldTypeDefaultDescription
orderstring[]–List of provider slugs to try in order (e.g. ["anthropic", "openai"]).

If you don't set this field, the router load balances across the top providers to maximize uptime. With order set, AnyRouter tries the listed providers one at a time and falls back to other providers if none are operational. If you don't want any other providers, disable fallbacks.

Example: Specifying providers with fallbacks

This example skips OpenAI (which doesn't host Mixtral), tries Together, and then falls back to the normal provider list:

curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer <ANYROUTER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "mistralai/mixtral-8x7b-instruct",
    "messages": [{"role": "user", "content": "Hello"}],
    "provider": { "order": ["openai", "together"] }
  }'

Example: Specifying providers with fallbacks disabled

With allow_fallbacks: false, the request fails if neither listed provider succeeds:

curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer <ANYROUTER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "mistralai/mixtral-8x7b-instruct",
    "messages": [{"role": "user", "content": "Hello"}],
    "provider": {
      "order": ["openai", "together"],
      "allow_fallbacks": false
    }
  }'

Targeting Specific Provider Endpoints

Each provider on AnyRouter may host multiple endpoints for the same model — a default endpoint and a specialized "turbo" endpoint, or region-specific endpoints like google-vertex/us-east5. Copy the exact provider slug from the provider list on the model detail page to target a specific endpoint.

Base Slug Matching

When you use a base provider slug (e.g. "google-vertex") in any provider routing field (order, only, or ignore), it matches all endpoints for that provider, including variants and regions. For example, "google-vertex" matches google-vertex, google-vertex/us-east5, google-vertex/us-central1, and so on.

To target a specific variant or region, use the full slug including the suffix (e.g. "google-vertex/us-east5" or "deepinfra/turbo").

Slug in requestWhat it matches
"google-vertex"All Google Vertex endpoints (every region)
"google-vertex/us-east5"Only the us-east5 region endpoint
"deepinfra"All DeepInfra endpoints (default + turbo)
"deepinfra/turbo"Only the DeepInfra turbo endpoint

Example: Targeting a specific endpoint variant

DeepInfra offers DeepSeek R1 through multiple endpoints — a default endpoint (deepinfra) and a turbo endpoint (deepinfra/turbo). Pin to the turbo endpoint:

curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer <ANYROUTER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek/deepseek-r1",
    "messages": [{"role": "user", "content": "Hello"}],
    "provider": {
      "order": ["deepinfra/turbo"],
      "allow_fallbacks": false
    }
  }'

This is especially useful when you want to consistently use a specific variant of a model from a particular provider.

To route to all endpoints of a provider (across all regions and variants), use the base slug without a suffix. For example, "google-vertex" routes across all Vertex AI regions.

Requiring Providers to Support All Parameters

Restrict requests only to providers that support all parameters in your request using require_parameters.

FieldTypeDefaultDescription
require_parametersbooleanfalseOnly use providers that support all parameters in your request.

With the default routing strategy, providers that don't support every LLM parameter in your request can still receive the request but will ignore unknown parameters. With require_parameters: true, such providers are skipped entirely.

Example: Excluding providers that don't support JSON formatting

curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer <ANYROUTER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-5.4-mini",
    "messages": [{"role": "user", "content": "Hello"}],
    "provider": { "require_parameters": true },
    "response_format": { "type": "json_object" }
  }'

Requiring Tool Calling and Context Size

Use these fields when anyrouter/auto (or another auto-routing model) should only pick members that can call tools or that have a large context window. The same fields persist on presets.

FieldTypeDescription
require_paramsstring[]Hard requirements. ["tools"] keeps only models that advertise tool calling. Distinct from boolean require_parameters.
min_contextnumberMinimum context window in tokens. 1_000_000 keeps only members with at least 1M context.

A non-empty tools array on the request also requires tool-calling members, even if you omit require_params.

curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer <ANYROUTER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anyrouter/auto",
    "messages": [{"role": "user", "content": "Plan the next step"}],
    "provider": {
      "sort": "exacto",
      "min_context": 1000000,
      "require_params": ["tools"]
    }
  }'

Save the same object on a preset and call @presets/<slug> — anyrouter/auto honors those floors on every request.

Requiring Providers to Comply with Data Policies

Restrict requests to providers that comply with your data policies using data_collection.

FieldTypeDefaultDescription
data_collection"allow" | "deny""allow"Control whether to use providers that may store data.
  • allow (default): allow providers which may store user data non-transiently and may train on it.
  • deny: use only providers which do not collect user data.

Some model providers may log prompts. AnyRouter flags these with a Data Policy tag on each model's provider list. This is not a definitive source for third-party data policies — it represents our best knowledge.

Example: Excluding providers that don't comply with data policies

curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer <ANYROUTER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-5.4-mini",
    "messages": [{"role": "user", "content": "Hello"}],
    "provider": { "data_collection": "deny" }
  }'

Zero Data Retention Enforcement

Enforce Zero Data Retention (ZDR) on a per-request basis using zdr, ensuring your request routes only to endpoints that do not retain prompts.

FieldTypeDefaultDescription
zdrboolean–Restrict routing to only ZDR endpoints.

When zdr is true, the request only routes to endpoints with a Zero Data Retention policy. When false or unset, it has no effect.

Example: Enforcing ZDR for a specific request

curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer <ANYROUTER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-5.4",
    "messages": [{"role": "user", "content": "Hello"}],
    "provider": { "zdr": true }
  }'

This is useful when you don't want to globally enforce ZDR but need to ensure specific requests only route to ZDR endpoints.

Distillable Text Enforcement

Restrict requests to models where the author has allowed text distillation, using enforce_distillable_text.

FieldTypeDefaultDescription
enforce_distillable_textboolean–Restrict routing to only models that allow text distillation.

Useful for applications that need to ensure their requests only use models where output can be used for training purposes, such as when building datasets for fine-tuning or distillation workflows.

curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer <ANYROUTER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/llama-3.3-70b-instruct",
    "messages": [{"role": "user", "content": "Hello"}],
    "provider": { "enforce_distillable_text": true }
  }'

Disabling Fallbacks

To guarantee your request is served only by the top (lowest-cost) provider, disable fallbacks. Often combined with order to restrict the candidate list to your chosen providers.

curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer <ANYROUTER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-5.4-mini",
    "messages": [{"role": "user", "content": "Hello"}],
    "provider": { "allow_fallbacks": false }
  }'

Allowing Only Specific Providers

Allow only specific providers for a request by setting only.

FieldTypeDefaultDescription
onlystring[]–List of provider slugs to allow for this request.

Allowing only some providers significantly reduces fallback options and limits request recovery.

Example: Allowing Azure for a GPT request

curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer <ANYROUTER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-5.4-mini",
    "messages": [{"role": "user", "content": "Hello"}],
    "provider": { "only": ["azure"] }
  }'

Ignoring Providers

Skip specific providers for a request by setting ignore.

FieldTypeDefaultDescription
ignorestring[]–List of provider slugs to skip for this request.

Ignoring multiple providers significantly reduces fallback options and limits request recovery.

Example: Ignoring DeepInfra for a Llama request

curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer <ANYROUTER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/llama-3.3-70b-instruct",
    "messages": [{"role": "user", "content": "Hello"}],
    "provider": { "ignore": ["deepinfra"] }
  }'

Quantization

Quantization reduces model size and computational requirements while aiming to preserve quality. Most LLMs today use FP16 or BF16 for training and inference, cutting memory requirements in half compared to FP32. Some deployments use FP8 or integer quantization to reduce size further (e.g., INT8, INT4).

FieldTypeDefaultDescription
quantizationsstring[]–List of quantization levels to filter by (e.g. ["int4", "int8"]).

Quantized models may exhibit degraded performance for certain prompts depending on the method used.

Quantization Levels

By default, requests are load-balanced across all available providers, ordered by price. To filter by quantization level, specify quantizations with one or more of:

  • int4: Integer (4 bit)
  • int8: Integer (8 bit)
  • fp4: Floating point (4 bit)
  • fp6: Floating point (6 bit)
  • fp8: Floating point (8 bit)
  • fp16: Floating point (16 bit)
  • bf16: Brain floating point (16 bit)
  • fp32: Floating point (32 bit)
  • unknown: Unknown

Example: Requesting FP8 quantization

curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer <ANYROUTER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta/llama-3.1-8b-instruct",
    "messages": [{"role": "user", "content": "Hello"}],
    "provider": { "quantizations": ["fp8"] }
  }'

Maximum Price

To filter providers by price, specify max_price with the highest provider pricing you will accept.

For example, {"prompt": 1, "completion": 2} routes only to providers with prompt token pricing ≤ $1/M and completion token pricing ≤ $2/M.

Some providers support per-request pricing — use the request field. image is also available for the max price per image input.

Practically, this is often combined with sort to express "use the provider with the highest throughput, as long as it doesn't cost more than $X/M tokens":

curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer <ANYROUTER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/llama-3.3-70b-instruct",
    "messages": [{"role": "user", "content": "Hello"}],
    "provider": {
      "sort": "throughput",
      "max_price": { "prompt": 1, "completion": 2 }
    }
  }'

Unlike preferred_min_throughput and preferred_max_latency, max_price is a hard filter — if no provider matches, the request fails with a 404 rather than falling back to a more expensive provider.

Provider-Specific Headers

Some upstreams support beta features that can be enabled through special headers. AnyRouter passes through a curated set of provider beta headers when present.

Anthropic Beta Features

When routing to Anthropic-backed models (Claude), you can request specific beta features with the anthropic-beta header. AnyRouter forwards supported values.

Supported Beta Features

FeatureHeader ValueDescription
Fine-Grained Tool Streamingfine-grained-tool-streaming-2025-05-14More granular streaming events during tool calls, with real-time updates as tool arguments are generated.
Interleaved Thinkinginterleaved-thinking-2025-05-14Allows Claude's thinking/reasoning to be interleaved with regular output rather than appearing as a single block.
Structured Outputsstructured-outputs-2025-11-13Enables strict tool use, validating tool parameters against your schema to ensure correctly-typed arguments.

AnyRouter manages some Anthropic beta features automatically:

  • Prompt caching and extended context are enabled based on model capabilities.
  • Structured outputs for JSON schema response format (response_format.type: "json_schema") — the header is applied automatically.

For strict tool use (strict: true on tools), you must pass anthropic-beta: structured-outputs-2025-11-13 explicitly. Without it, AnyRouter strips the strict field and routes normally.

Example: Enabling Fine-Grained Tool Streaming

curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer <ANYROUTER_API_KEY>" \
  -H "Content-Type: application/json" \
  -H "anthropic-beta: fine-grained-tool-streaming-2025-05-14" \
  -d '{
    "model": "anthropic/claude-sonnet-4.6",
    "messages": [{"role": "user", "content": "What is the weather in Tokyo?"}],
    "tools": [{
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Get the current weather for a location",
        "parameters": {
          "type": "object",
          "properties": { "location": { "type": "string" } },
          "required": ["location"]
        }
      }
    }],
    "stream": true
  }'

Example: Enabling Interleaved Thinking

curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer <ANYROUTER_API_KEY>" \
  -H "Content-Type: application/json" \
  -H "anthropic-beta: interleaved-thinking-2025-05-14" \
  -d '{
    "model": "anthropic/claude-sonnet-4.6",
    "messages": [{"role": "user", "content": "Solve this step by step: What is 15% of 240?"}],
    "stream": true
  }'

Combining Multiple Beta Features

Separate values with commas:

anthropic-beta: fine-grained-tool-streaming-2025-05-14,interleaved-thinking-2025-05-14

Beta features are experimental and may change or be deprecated by Anthropic. Check Anthropic's documentation for the latest list of available beta features.

Terms of Service

You can view the terms of service for each upstream provider on the Providers page. You may not violate the terms of service or policies of third-party providers that power the models on AnyRouter.

Verify

Send a request with a provider block and confirm it is honored — for example, pin a single provider with fallbacks off:

curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer sk-ar-your-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-5.4-mini",
    "messages": [{"role": "user", "content": "Hello"}],
    "provider": { "only": ["openai"], "allow_fallbacks": false }
  }'

If the pinned provider is unavailable and fallbacks are off, the request fails loudly — that is the expected, non-silent behavior.

Troubleshooting

My request fails with a 404 when I set max_price

max_price is a hard filter. If no upstream is within your ceiling, the request fails with a 404 rather than routing to a more expensive provider. Raise the cap or remove it.

A performance threshold isn't excluding slow providers

preferred_min_throughput and preferred_max_latency are preferences, not hard filters — providers that miss them are deprioritized to the fallback pool, not excluded. Use only/ignore to hard-exclude a provider.

My provider slug matches more endpoints than expected

A base slug like "google-vertex" matches every region and variant. Use the full slug with its suffix ("google-vertex/us-east5", "deepinfra/turbo") to target one endpoint.

Related