Skip to content
Checking status

Unify every model one gateway.

One OpenAI-compatible API, CLI and MCP gateway in front of every provider — and every other router. Free models, $4 credit every month on Go, and your own keys at no markup.

Free to start — $4 credit every month on Go, no card if you donate a key. Unlock the shared pool for $2/mo or by donating one working key.

Use it withOpenAIxAIAnthropic+ 206 models across 17 providers
anyrouter ~ openai (python)
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://anyrouter.dev/api/v1",
    api_key=os.environ["ANYROUTER_API_KEY"],
)

resp = client.chat.completions.create(
    model="anthropic/claude-sonnet-4.6",
    messages=[{"role": "user", "content": "Hi"}],
)
ANYROUTER_API_KEYget key

Crowdsourced free tokens

We pool everyone's free & trial keys into one big shared pool — free tokens for everyone. Donate a key, grow the pool, ride free.

01

One endpoint. Every model.

Keep your SDK. One upstream rate-limits — your request doesn't notice.

Point your SDK at one base URL and reference models as provider/model. Every call then runs the same pipeline: authenticate, pick an upstream, retry the next one on failure, cache what repeats, and log every attempt.

Between your call and the model · per request

Key balancing

Spreads load across your keys and quarantines burned ones automatically — no client changes.

Automatic failover

Retries and falls back across providers and pooled keys the instant one errors or rate-limits.

Prompt caching

Reuses cached context across calls to cut repeat token cost and time-to-first-token.

Per-request logs

Every fallback hop, token count and cost captured per attempt — traced and queryable.

03

Everything in one gateway

One key. Every capability you'd otherwise wire up yourself.

No add-ons, no tiers gating the basics. Every capability below ships on the same base URL and the same API key.

04

Why teams switch

Three things that are structurally hard to copy.

Never taxed

Bring your own keys, keep every cent.

Route through your own provider keys and AnyRouter takes no cut — no deposit fee, no per-token markup, no cap. What you pay upstream is exactly what you pay.

Observable

Debug any request.

Every fallback hop, status, latency, token count and cost — captured per attempt. A “debug this request” trace nobody else ships at this tier.

Portable

Your config follows you.

Keys, presets and skills live in AnyRouter, inject locally on demand, and wipe clean on exit. Same setup on your laptop, a server, CI, or a teammate's box.

05

Free models, funded by everyone

Shared key pool

Every member signs up for a provider's free/trial tier — NVIDIA NIM, the Gemini free tier, the Groq free tier, and more — and donates that key. The quotas add up into one pool that every member — you included — can call. Donating a key makes your Go plan free.

06

Unlock Go — two doors, same room

How to join

AnyRouter runs on Go — $2/mo, or donate one free-tier provider key (donors ride free). Every Go comes with $4 credit each month.

Create your account

Sign up in seconds — no card required. You land on the dashboard with your first API key ready to copy.

07

The cheapest way in

Three ways to start

A $4 monthly credit and free models come with Go — $2/mo, or free when you donate a provider key. Or bring your own keys at no markup, no card required.

Go plan

$4 credit / month

On Go — $2/mo or a donated key. Here's how far $4 goes on fast models.

Tokens for $4, blended 75% input / 25% output.

Go plan

Free models, 1000/day

Route via anyrouter/free on Go — $0 per token, up to 1000 requests/day.

Apple Foundation Model (on-device)
$0 in · $0 out
Dots3-Note Preview
$0 in · $0 out
Gemma 4 31B
$0 in · $0 out
Ling-3.0-flash-Fin
$0 in · $0 out
Ling-3.0-flash-Sante
$0 in · $0 out
Ling-3.0-flash-VL
$0 in · $0 out
Ling-3.0-tiny
$0 in · $0 out
LFM 2.5 2.6B
$0 in · $0 out
Nex-N2.5-Mini
$0 in · $0 out
Nex-N2.5-Pro
$0 in · $0 out
Nemotron-3 Nano 30B
$0 in · $0 out
Nemotron-3 Super 120B
$0 in · $0 out
Browse free models
Bring your own keys

Your keys, no markup

12 providers ship a free tier — billed by the provider only, never by us.

OpenCode Zen
Google AI Studio
NVIDIA NIM
Ollama Cloud
OpenRouter
Hugging Face
Add your keys
08

Sign in with AnyRouter

Add AI to your app — your users bring their own AnyRouter

One OAuth button hands your app a scoped, temporary key per user. Every request is billed to that user's own account — no API keys for you to collect, store, or pay for.

Your app's sign-in screen

Click it — see what your users see.

  • OAuth 2.1 + PKCE

    Standard authorization-code flow, public client, no secret to leak. Works from a browser-only app.

  • Inference-only scoped key

    The token can run AI requests and read a basic profile — never keys, billing, or account settings.

  • Per-user billing & revocation

    Each user pays from their own credits or free tier, and can revoke your app any time from their dashboard.

09

206+ models · 17 providers · September 29, 2026

One growing catalog, always current

StepFun's next-generation flagship base model joins the catalog BYOK-only on AIHubMix — 1M-token context, 64K max output, text/image/video input, and adjustable reasoning at $1.08 / $3.0888 per million input / output tokens.

1M

Anthropic's Claude Sonnet 5.5 is the successor to Claude Sonnet 5 — the best combination of speed and intelligence in the lineup. It is the direct upgrade for well-scoped everyday work (building features, fixing bugs, drafting and review) at the same $2 / $10 per million input / output rates, with adaptive thinking on by default at `high` effort, 1M-token context, 128K max output, vision input, and tool use. Faster than Opus 5.5, cheaper, same window.

1M

Step 5 Preview is StepFun's next-generation flagship base model designed for real-world tasks, with a focus on programming and professional expertise. A multimodal (text, image, video in / text out) base model with a 1M-token context and adjustable reasoning effort.

8K

Gemini Embedding 2 is Google's first multimodal embedding model, mapping text, images, video, audio, and PDFs into one 3,072-dimension vector space for cross-modal semantic search, document retrieval, and recommendations over 100+ languages. Upstream accepts 8,192 input tokens and MRL-truncates to 128–3,072 dimensions. AnyRouter exposes the text path only — the OpenAI-compatible /embeddings body is a text `input`, so the extra upstream modalities are deliberately not declared here.

33K

Qwen3 Embedding 4B is the 4B size in the Qwen3 embedding and reranking family, between the 0.6B and 8B cuts. It embeds text over a 32,768-token window at 2,560 dimensions, supports 100+ languages, is instruction-aware for task prefixes, and supports Matryoshka truncation from 32 to 2,560 dimensions. It ranks below the 8B cut on MTEB multilingual but well above the 0.6B cut, at a much lower cost.

Span-01 Lite is a behavior-scoring classifier from Respan — the free, lighter tier of Span-01, a 4B model trained with RLAIF on classification reasoning rather than fixed taxonomies. Send a conversation span and a set of plain-language behavior definitions and get a present / absent / not_observable probability per behavior in a single forward pass, with no token-by-token generation. Respan reports it ahead of Jev 1.13 and Sonnet 5 on their English and multilingual behavior-detection suites. Call POST /api/v1/decisions (POST /api/v1/systemone is the same handler).

$2.00 in
$6.00 out
1M

Fugu Max is the largest tier of Sakana AI's Fugu line — a 1M-token-context model that takes text, images, and files and returns text. It joins sakana/fugu and sakana/fugu-ultra as a distinct tier, and is not a v2 of fugu-ultra (that one is folded into the existing row as an alias, not a new listing).

8K

Hy-MT2 30B A3B is Tencent's translation model — a 30B-parameter MoE with 3B active parameters, tuned for machine translation with a short 8,192-token window. It is the first translation family in the AnyRouter catalog and the largest of the three published sizes (1.8B, 7B, 30B-A3B); only the 30B-A3B cut is listed here.

524K

Solar Mini 4 is Upstage's small, fast tier in the Solar 4 generation: a text-only model with a 524,288-token context window at roughly $0.05 / $0.20 per million input / output tokens. It is the cheap sibling of upstage/solar-pro4, not a packaging of it, so it gets its own listing row.

1M

GLM 5.3 Prime is the top reasoning tier of Z.ai's GLM 5.3 line, above the Flash and FlashX cuts, with a 1M-token context window. It sits alongside z-ai/glm-5.3 and z-ai/glm-5.3-flash as a separately priced official variant rather than a packaging of them, so it keeps its own listing row.

$0.01 in
$0 out
8K

BAAI's BGE-M3 is a multilingual text embedding model that handles dense, sparse, and multi-vector retrieval in one model, across 100+ languages and an 8,192-token input window. It returns a 1,024-dimension dense vector.

8K

BAAI's bge-multilingual-gemma2 is a compact multilingual text embedding model built on Gemma 2, trained across 100+ languages with an 8,192-token input window. It returns a 3,584-dimension dense vector.

131K

Mistral Small 3.2 is a 24B instruction model with a 131k context window, improved instruction following and function calling over Mistral Small 3.1, and the same Apache 2.0 license.

1.1M

MiMo V2.6 Flash is Xiaomi's full-modality reasoning model for high-frequency invocations and large-scale professional workloads. It keeps the 1M-token context window and up to 128K output tokens of MiMo V2.6 Pro at roughly a third of the input price, making it the best-balanced pick when reasoning quality matters more than peak capability.

1.1M

MiMo V2.6 Pro is Xiaomi's trillion-parameter, natively omni-modal flagship reasoning model for complex projects, long-horizon agentic tasks, coding, research, and professional knowledge work. It supports a 1M-token context window with up to 128K output tokens, and keeps the V2.5 API price while improving reasoning and long-task completion.

8K

Kev 4B is a small open-weight decision model from Jared Palmer — a LoRA adapter and pointer head on Qwen3.5-4B-Base. Send state and typed questions (noul / choice / score) and get a calibrated probability per question in one forward pass, not generated text. Call POST /api/v1/decisions (POST /api/v1/systemone is the same handler).

32K

Fastino GLiNER2.5-Decide is a decision model: send application state and typed questions (noul / choice / score) and get structured answers with confidence — not generated chat text. Call POST /api/v1/decisions (POST /api/v1/systemone is the same handler). Upstream, GLiNER2.5 runs schema-driven classifications on an OpenAI-compatible chat-completions API; the gateway adapts the System One envelope for this hop.

1M

LongCat-2.5-Preview is Meituan's sparse mixture-of-experts frontier model — ~1.6T total parameters with ~48B active — over a 1M-token context window. It adds native multimodal understanding (image input), stronger coding ability, and a `thinking` toggle, and is built for long-horizon agentic work across terminals, browsers, GUIs, spreadsheets, and design tools. Served over the OpenAI-compatible and Anthropic-compatible APIs.

260K

Inception Labs' Mercury 2.5 is a diffusion language model for fast text generation and reasoning, with a 260,000-token context window.

164KZDR

Llama Guard 4 is a 12B natively multimodal safety classifier for content safety classification on LLM inputs and responses. It acts as an LLM itself: it generates text indicating whether a prompt or response is safe or unsafe and, if unsafe, lists the violated categories (MLCommons hazards taxonomy S1–S14). Use it as a guardrail/judge hop in front of or behind chat models.

1.1M

OpenAI's GPT-6 Luna is the efficient GPT-6 tier for focused, high-throughput tasks, with configurable reasoning, tool use, image input, and a 1.05M-token context window.

$2.00 in
$10.00 out
1.1M

OpenAI's GPT-6 Sol is built for complex coding and agentic workflows, with configurable reasoning, tool use, image input, and a 1.05M-token context window. It is the high-end GPT-6 tier below GPT-6 Astra.

131K

OpenAI's open-weight safety reasoning model built on gpt-oss-20b (Apache 2.0). A 21B-parameter MoE tuned for safety tasks: content classification, LLM input/output filtering, and trust & safety judgments with reasoning traces. Use it as a guardrail/judge hop in front of or behind chat models.

262K

Qwen3-235B-A22B-Instruct-2507 is Qwen's updated non-thinking MoE model for multilingual instruction following, coding, tool use, and long-context work. It activates 22B of 235B parameters and provides a native 262K context window.

131K

L3.1 Euryale 70B is Sao10k's text-generation model for creative roleplay, with a 131K context window. It is the successor to Euryale L3 70B v2.1.

10

FAQ

Frequently asked questions

What is AnyRouter?

One OpenAI-compatible API, CLI and MCP gateway in front of every provider — and every other router. Free models, $4 credit every month on Go, and your own keys at no markup.

How do I call models through AnyRouter?

Point your SDK at one base URL and reference models as provider/model. Every call then runs the same pipeline: authenticate, pick an upstream, retry the next one on failure, cache what repeats, and log every attempt.

Do you take a cut if I bring my own keys?

Route through your own provider keys and AnyRouter takes no cut — no deposit fee, no per-token markup, no cap. What you pay upstream is exactly what you pay.

How do I get started?

Sign up in seconds — no card required. Go is the plan that switches AnyRouter on: pay $2 / month, or donate 1 free key. Go comes with $4 credit every month, unlimited free models, and access to the shared key pool.