Chat Completions API
POST /api/v1/chat/completions — the OpenAI-compatible chat endpoint. Send messages, get a completion, across every model in the catalog.
Send a sequence of messages to a model and receive a completion. This is the primary AnyRouter inference endpoint and a drop-in replacement for OpenAI's POST /v1/chat/completions — point your OpenAI client at the AnyRouter base URL and it works unchanged.
/api/v1/chat/completionsAuthenticate with an LLM API key (sk-ar-v1-…) from /dashboard/keys. Management keys (ak_…) are not accepted on inference routes.
Request
| Field | Type | Required | Description |
|---|---|---|---|
model | string | yes | provider/model id (e.g. openai/gpt-5.4-mini). See List models. |
messages | array | yes | Ordered conversation. Each message has role (system / user / assistant / tool) and content. |
temperature | number | no | 0–2. Lower is more deterministic. Default: 1. |
top_p | number | no | 0–1. Nucleus sampling. Default: 1. |
max_tokens | integer | no | Upper bound on output tokens. |
stream | boolean | no | If true, returns a streaming SSE response. |
stop | string | string[] | no | Stop sequences. |
tools | array | no | Function-calling / tool-use spec. |
tool_choice | string | object | no | Force a specific tool or auto. |
response_format | object | no | { "type": "json_object" } for guaranteed JSON. |
seed | integer | no | Deterministic sampling seed (provider-dependent). |
user | string | no | Stable end-user identifier for abuse monitoring. |
session_id | string | no | Sticky-session id (≤256 chars) that groups related requests in the Request Logs dashboard. Also settable via the x-session-id header; the body field wins if both are sent. |
provider | object | no | Request-level routing preferences such as only, ignore, order, sort, allow_fallbacks, and max_price. See Smart Routing. |
trace | object | no | Optional tracing metadata. Supports trace_id, trace_name, span_name, generation_name, and parent_span_id. |
You can also force a provider for a single request with X-AnyRouter-Provider: openai, or a comma-separated fallback list such as X-AnyRouter-Provider: groq,openai. Every response returns X-Request-ID for support correlation and X-AnyRouter-Provider for the upstream that served it.
Routing preferences
provider biases or constrains upstream selection without changing the model id:
{
"provider": {
"only": ["OpenAI", "Groq"],
"order": ["Groq", "OpenAI"],
"sort": "ttft",
"allow_fallbacks": true,
"max_price": {
"prompt": "1.00",
"completion": "2.00"
}
}
}
Use it when you want a lower-latency provider, a strict price ceiling, or a reproducible preferred backend order. AnyRouter tracks live TTFT and TPS per upstream from real traffic and re-ranks every request.
provider.sort | Sorts by | Direction | When to use |
|---|---|---|---|
cost | Listed input price per million tokens | Lowest price first | High-volume, cost-sensitive work |
ttft | Median time-to-first-token (ms) | Lowest latency first | Latency-sensitive workloads |
tps | Median tokens-per-second throughput | Highest first | Long-output generation |
Shortcut suffixes on the model id pick the sort for you: model:floor → cost, model:fast → ttft, model:nitro → tps. Legacy names price, latency, and throughput are accepted aliases for cost, ttft, and tps. Supported routing fields are only, ignore, order, sort, allow_fallbacks, max_price.prompt, max_price.completion, preferred_max_latency, and preferred_min_throughput; other fields are accepted for client compatibility but do not affect local routing.
Response
{
"id": "chatcmpl-8yWq4JqLfEjJ9L",
"object": "chat.completion",
"created": 1760000000,
"model": "anthropic/claude-sonnet-4.6",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "NAT maps many private IP addresses to one public IP so devices behind a router can share a single internet connection."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 28,
"completion_tokens": 30,
"total_tokens": 58
}
}
finish_reason | Meaning |
|---|---|
stop | Model completed naturally or hit a stop sequence. |
length | Hit max_tokens. |
tool_calls | Model invoked one or more tools; caller must run them and resume. |
content_filter | Upstream safety filter triggered. |
error | Upstream provider error; see error on the response. |
Examples
curl https://anyrouter.dev/api/v1/chat/completions \
-H "Authorization: Bearer sk-ar-your-key" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-4.6",
"messages": [
{ "role": "system", "content": "You are a concise assistant." },
{ "role": "user", "content": "Explain NAT in one sentence." }
],
"max_tokens": 200
}'
from openai import OpenAI
client = OpenAI(
base_url="https://anyrouter.dev/api/v1",
api_key="sk-ar-your-key",
)
response = client.chat.completions.create(
model="anthropic/claude-sonnet-4.6",
messages=[
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "Explain NAT in one sentence."},
],
max_tokens=200,
)
print(response.choices[0].message.content)
import OpenAI from "openai"
const openai = new OpenAI({
baseURL: "https://anyrouter.dev/api/v1",
apiKey: "sk-ar-your-key",
})
const completion = await openai.chat.completions.create({
model: "anthropic/claude-sonnet-4.6",
messages: [
{ role: "system", content: "You are a concise assistant." },
{ role: "user", content: "Explain NAT in one sentence." },
],
max_tokens: 200,
})
console.log(completion.choices[0].message.content)
Tool use
Pass a tools array with function schemas, then handle any returned tool_calls:
curl https://anyrouter.dev/api/v1/chat/completions \
-H "Authorization: Bearer sk-ar-your-key" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-5.4-mini",
"messages": [
{ "role": "user", "content": "What is the weather in Tokyo?" }
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city",
"parameters": {
"type": "object",
"properties": { "city": { "type": "string" } },
"required": ["city"]
}
}
}
]
}'
const response = await client.chat.completions.create({
model: "openai/gpt-5.4-mini",
messages: [{ role: "user", content: "What is the weather in Tokyo?" }],
tools: [
{
type: "function",
function: {
name: "get_weather",
description: "Get the current weather for a city",
parameters: {
type: "object",
properties: { city: { type: "string" } },
required: ["city"],
},
},
},
],
})
See Tool use for a deeper walkthrough.
Errors
All errors follow the standard AnyRouter JSON envelope with a stable error.code.
| Status | Meaning | Fix |
|---|---|---|
| 400 | Malformed body or invalid parameters. | Check the request against the field table above. |
| 401 | Missing or invalid API key. | Send a valid sk-ar-… key as Authorization: Bearer. |
| 402 | Out of credits or upstream BYOK billing failure. | Top up credits, or read metadata.upstream_message. |
| 404 | Unknown model. | Verify the id with List models. |
| 429 | Rate limit hit. | Back off; inspect X-RateLimit-* response headers. |
| 502 | Upstream 5xx or every fallback failed. | Retry; see metadata.upstream_attempts. |
See Errors for the full status table and retry recipes.
