Free-tier Aggregation
Combine multiple providers' free API tiers behind one endpoint, with smart key balancing, automatic quota detection, and an opt-in paid fallback.
Every provider hands out a free tier, but each has its own key, endpoint, and quota — so squeezing real capacity out of them means juggling a drawer of keys by hand. Free-tier aggregation collapses that into one OpenAI-compatible endpoint: bring your free keys from Groq, Cerebras, SambaNova, NVIDIA, Mistral, OpenRouter, GitHub Models, Cohere, Z.ai, and more, and each request routes to a key that still has budget — with automatic cooldown when a provider rate-limits you and a clean fall-through when one is exhausted.
Overview
You combine provider branches into a preset (a small routing tree) and call it like any other model:
curl https://anyrouter.dev/api/v1/chat/completions \
-H "Authorization: Bearer sk-ar-your-key" \
-H "Content-Type: application/json" \
-d '{
"model": "@presets/duyet-coding",
"messages": [{"role": "user", "content": "Hello"}]
}'
Embeddings work the same way:
curl https://anyrouter.dev/api/v1/embeddings \
-H "Authorization: Bearer sk-ar-your-key" \
-H "Content-Type: application/json" \
-d '{ "model": "@presets/my-embeddings", "input": "..." }'
The preset's kind must match the endpoint (chat vs embedding). Mismatches return 400 preset_kind_mismatch.
How it works
Key balancing and cooldown
When you send model: "@presets/<your-slug>" to /api/v1/chat/completions or /api/v1/embeddings, the resolver picks a branch, picks a key inside that branch using your per-provider key strategy (round-robin, weighted, fallback, or random — set on the BYOK page), and dispatches the request.
If the chosen key returns 429, it's placed on a 120-second cooldown (or longer if the provider sends a Retry-After header) and the next request automatically tries the next key. If every key in every branch is exhausted, the request gets a 429 — unless you've turned on Paid fallback.
Paid fallback
Each preset has a Paid fallback switch. When off (default), exhaustion returns 429. When on, the preset appends a synthetic "anyrouter platform" branch at the lowest priority — when every BYOK branch is unavailable, the request is served on platform credits instead. Costs land on the same cost_usd column you already see in the dashboard.
Reading the dashboard
The Capacity card at the top of every preset page shows three independent meters:
| Meter | What it measures |
|---|---|
| Tokens | Sum of daily token budget − tokens used today across every BYOK key in the preset |
| Requests | Same shape for request counts |
| Platform $ | Spend vs. budget when the preset has a paid-fallback branch |
A status row collapses every key into four dots:
●healthy◐cooldown (Retry-Afterseconds remaining)○exhausted (≥95% of daily budget used)✕invalid (provider rejected the key on the last probe)
A countdown row shows the next reset for each window kind, so you know when capacity comes back.
Re-probing
Limits change — providers raise free tiers, you upgrade your plan, your weekly window rolls over. Click Probe all keys on the dashboard to re-run every probe concurrently. Each probe has a 3-second budget so the page stays responsive; results land on D1 and the meters refresh on the next view.
You can also re-probe a single key via POST /api/v1/byok/probe with { "key_id": "byok_..." }.
Configure
Add your provider keys
Open Dashboard → BYOK, click Add key, choose a provider, paste the API key, give it a label, and click Validate & Add. AnyRouter validates the key against the provider and shows detected limits inline:
- Z.ai — 5-hour token window plus per-minute request rate
- Anthropic (Claude OAuth) — 5-hour, 7-day, and per-model windows
- DeepSeek — USD credit balance
- OpenRouter — USD credit balance and key-level rate limit
- OpenAI — daily request and token counts (admin keys only)
- Groq / Cerebras / SambaNova / NVIDIA / Mistral / Cohere / GitHub Models — per-minute request and token budget read from response headers
If a provider doesn't expose usage data, the key is saved as Unvalidated with an amber badge. You can fill in limits manually under Advanced on the same form — but you don't have to.
Build a preset
A preset is a routing tree — Preset → [Branch 1, Branch 2, …] → BYOK keys (auto-derived). Each branch is a (model, upstream) pair; the same model can appear twice via different upstreams (for example z-ai/glm-4.7-flash via your Z.ai key in one branch and the same AnyRouter model id via your OpenRouter key in another).
Open Dashboard → Presets → New. For each branch: pick a model, pick the upstream (for example Z.ai Standard API, Z.ai Coding Plan, OpenRouter, or AnyRouter platform credits), and set a weight (used in weighted mode) or priority (used in fallback mode).
Switch modes on the root:
- Fallback — try Branch 1 first, only move to Branch 2 if every key in Branch 1 is cooled down or exhausted.
- Weighted — distribute requests across branches by weight (e.g. 70/20/10).
Call the preset
Send model: "@presets/<your-slug>" to /api/v1/chat/completions or /api/v1/embeddings, and the resolver handles branch and key selection for you.
What's not included
A few capabilities are deliberately out of scope:
- OAuth-subscription tiers (Claude Code, GitHub Copilot, Cursor) — they're typically tied to subscriber terms that don't allow third-party intermediation.
- Sticky sessions — every turn re-resolves the preset; if you need turn-to-turn affinity, pin a single model in your request instead.
- Background probing — probes run when you open the dashboard, not on a cron.