API Overview
Choose the right AnyRouter inference surface — Chat Completions, Messages, Responses, Embeddings — plus catalog and status endpoints.
AnyRouter exposes four core inference surfaces — Chat Completions, Messages, Responses, and Embeddings — plus catalog and account endpoints. Most OpenAI-compatible tooling works by overriding the base URL and API key; Anthropic-compatible tooling uses the Messages endpoint.
The full OpenAPI spec is browsable at /docs/api — every endpoint with live request/response samples and a "try it" runner.
Base URL
https://anyrouter.dev/api/v1
OpenAI-compatible clients that set the base URL to https://anyrouter.dev (and append /v1/... themselves) also work: /v1/models and /v1/chat/completions are the same API as /api/v1/....
All endpoints are versioned under /v1. Breaking changes bump the version (/v2) without removing older versions for a minimum 12-month deprecation window.
Choose an endpoint
Chat Completions — OpenAI-compatible chat, streaming, tool calls, and broad cross-provider routing.
Messages — Anthropic-compatible clients such as Claude Code or the Anthropic SDK.
Responses — modern agent and tool workflows that need structured output items.
Embeddings — semantic search, clustering, and classification.
Decisions — TypeSafe typed decisions (noul / choice / score), not chat.
OCR — extract text from images and documents as Markdown.
Models — discovery, context windows, capabilities, pricing, and provider routes.
Health — live status, checked_at, and per-component availability.
Authentication: which key for which endpoint
AnyRouter uses two distinct API key families. Pick the right one for the endpoint you're calling.
| Key family | Prefix | Use for | Get one at |
|---|---|---|---|
| LLM API key | sk-ar-v1- | Inference: /chat/completions, /messages, /responses, /embeddings, /decisions (/systemone is the same handler). Also reads /models, /generation/:id, /auth/key. | /dashboard/keys |
| Management API key | ak_ | Account administration: /keys (LLM-key CRUD), /presets, /auth/byok, /credits, /management-keys. | /dashboard/management-keys |
The two families do not cross over. Sending an ak_ key to /chat/completions returns 401; sending a sk-ar-v1- key to /api/v1/keys returns 401 with a hint pointing you here.
flowchart LR
App["Your app"]
App -->|"sk-ar-…"| LLM["LLM API key"]
App -->|"ak_…"| Mgmt["Management API key"]
LLM --> Inference["Inference<br/>POST /chat/completions<br/>POST /messages<br/>POST /responses<br/>POST /embeddings<br/>POST /decisions"]
LLM --> Read["Read-only<br/>GET /models<br/>GET /generation/:id<br/>GET /auth/key"]
Mgmt --> Admin["Account admin<br/>/keys (CRUD)<br/>/management-keys<br/>/presets<br/>/auth/byok<br/>/credits"]
classDef key fill:#fef3c7,stroke:#f59e0b,stroke-width:1px,color:#92400e
classDef endpoint fill:#f1f5f9,stroke:#cbd5e1,color:#0f172a
class LLM,Mgmt key
class Inference,Read,Admin endpoint
Quick decision rule:
- I want a model to answer a prompt →
sk-ar-LLM key. - I want to create, rotate, or revoke keys / manage BYOK / read balances →
ak_management key.
See Management API Keys for the full scope catalog and management lifecycle.
Common headers
| Header | Required | Description |
|---|---|---|
Authorization: Bearer <key> | Yes | LLM key (sk-ar-…) on inference routes; management key (ak_…) on account routes. |
Content-Type: application/json | POST | Required on POST endpoints. |
x-api-key | Messages | Anthropic-compatible auth header for /messages (accepts sk-ar-…). |
anthropic-version | Messages | Required by Anthropic-compatible Messages clients. |
X-AnyRouter-Provider | No | Force routing to a specific upstream provider. |
X-AnyRouter-Trace-Id | Response | Canonical trace id on every inference response. Paste it into the dashboard logs "Request ID" filter to open the exact call. |
X-Request-ID | Response | Legacy alias of X-AnyRouter-Trace-Id — same value, kept for back-compat. |
X-AnyRouter-Model | Response | Public catalog model id that served the request (including the resolved model for a virtual preset). |
X-AnyRouter-Upstream-Model | Response | Wire model name sent to the selected upstream. |
X-AnyRouter-Provider-Chain | Response | Ordered backend ids tried, joined with ->; contains route ids only, never credentials. |
X-AnyRouter-Credential-Source | Response | platform, own_byok, or pool for the winning attempt. This is a class label, not key material. |
X-AnyRouter-Credits-Balance-At-Admission | Response, when read | Workspace balance observed during paid-inference admission, before the asynchronous debit. It is omitted when the balance gate is skipped (free/BYOK-only). |
X-AnyRouter-Credits-Balance | GET /credits response | Current authoritative workspace balance returned by the credits read endpoint. |
X-RateLimit-* / RateLimit-* | Response | Authoritative free-tier daily limit, remaining count, UTC reset, and policy on admitted free-model requests; the free-tier daily values take precedence for that response. |
Endpoint summary
| Method | Path | Description |
|---|---|---|
| POST | /chat/completions | Chat completions (OpenAI-compatible) |
| POST | /ocr/extract | OCR — extract text as Markdown |
| POST | /messages | Anthropic Messages passthrough |
| POST | /responses | Responses for structured output and tool workflows |
| POST | /embeddings | Embeddings for search and clustering |
| POST | /decisions | Decisions TypeSafe decisions (not chat). Preferred path. |
| POST | /systemone | Same handler as /decisions |
| GET | /models | List models |
| GET | /models/:id | Model metadata |
| GET | /models/count | Total model count |
| GET | /generation/:id | Per-request stats (tokens, cost, latency) |
| GET | /providers | List upstream providers |
| GET | /credits | Remaining credit balance |
| GET | /auth/key | Current API key info |
| GET | /health | Gateway + provider status |
Error shape
All errors follow a consistent JSON envelope:
{
"error": {
"code": "payment_required",
"message": "Upstream provider \"cloudflare\" returned 402 — insufficient quota",
"metadata": {
"type": "upstream_error",
"upstream_backend": "cloudflare",
"upstream_status": 402,
"upstream_message": "You exceeded your current quota, please check your plan and billing details."
}
}
}
error.code is a stable machine-readable slug. The canonical set:
| HTTP | error.code | Notes |
|---|---|---|
| 400 | bad_request | Malformed body or invalid parameters. |
| 401 | unauthorized | Missing or invalid API key. |
| 402 | insufficient_balance | Out of credits on a paid model, or a held signup bonus still pending (signup_bonus_pending). Top up at Credits. Upstream BYOK billing failures may also surface as 402 with metadata.type: "upstream_error" — read metadata.upstream_message. |
| 403 | forbidden | Key lacks permission for the resource. |
| 404 | not_found | Unknown model or missing resource. |
| 408 | request_timeout | Upstream did not respond in time. |
| 413 | content_too_large | Request body exceeds the configured limit. |
| 422 | unprocessable_entity | Validation failed semantically. |
| 429 | rate_limit_exceeded / ip_rate_limit_exceeded / free_tier_daily_limit_exceeded | Rate limit hit. Honor Retry-After. X-RateLimit-Remaining / X-RateLimit-Reset appear only on the daily free-tier and Go 5-hour window limiters — not on the per-minute tier. |
| 500 | internal_server_error | Unexpected AnyRouter failure. |
| 502 | bad_gateway / upstream_exhausted | Upstream 5xx or every fallback failed; see metadata.upstream_attempts. |
| 503 | service_unavailable | Upstream temporarily unavailable. |
When AI Gateway BYOK rejects a call, AnyRouter parses the provider's verbatim message into error.metadata.upstream_message — that is the authoritative reason, not Cloudflare's wrapping text. See Errors for retry recipes and /docs/api for the full interactive reference.
Rate limits
Rate limits are enforced per-key. Every response carries at least:
X-RateLimit-Limit: <requests per minute>
X-RateLimit-Window: 60
X-RateLimit-Tier: <tier name>
RateLimit-Limit: <same as X-RateLimit-Limit>
RateLimit-Policy: <limit>;w=60
The per-minute tier does not emit X-RateLimit-Remaining or X-RateLimit-Reset — the edge limiter only reports the cap. The daily free-tier cap and Go's 5-hour window do emit Remaining / Reset (and matching RateLimit-* headers). See Rate Limits.
Tracing
POST /chat/completions and POST /responses also accept an optional request-body trace object:
{
"trace": {
"trace_id": "4f8c2b9670b44b49a2e71e7fd5c0b1d3",
"trace_name": "checkout-flow",
"span_name": "summarize-order"
}
}
AnyRouter uses this metadata to correlate internal request steps such as identity resolution, rate-limit checks, upstream selection, and upstream dispatch. If you do not send a trace object, AnyRouter still returns an X-Request-ID header you can use for support.
Versioning
Changes are categorized as:
- Additive (new fields, new endpoints) — ship continuously, no notice.
- Deprecation — announced with a 12-month sunset window and a
Deprecationresponse header. - Breaking — land on a new
/v2path;/v1remains available for the deprecation window.
Related
- Migrate from OpenAI — swap base URL and key
- Migrate from Anthropic — point Anthropic clients at Messages
- Migrate to Responses — adopt the structured lifecycle
- Errors — status codes and retry guidance