Skip to content

Free-tier Aggregation

Combine multiple providers' free API tiers behind one endpoint, with smart key balancing, automatic quota detection, and an opt-in paid fallback.

Every provider hands out a free tier, but each has its own key, endpoint, and quota — so squeezing real capacity out of them means juggling a drawer of keys by hand. Free-tier aggregation collapses that into one OpenAI-compatible endpoint: bring your free keys from Groq, Cerebras, SambaNova, NVIDIA, Mistral, OpenRouter, GitHub Models, Cohere, Z.ai, and more, and each request routes to a key that still has budget — with automatic cooldown when a provider rate-limits you and a clean fall-through when one is exhausted.

Overview

You combine provider branches into a preset (a small routing tree) and call it like any other model:

curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer sk-ar-your-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "@presets/duyet-coding",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

Embeddings work the same way:

curl https://anyrouter.dev/api/v1/embeddings \
  -H "Authorization: Bearer sk-ar-your-key" \
  -H "Content-Type: application/json" \
  -d '{ "model": "@presets/my-embeddings", "input": "..." }'

The preset's kind must match the endpoint (chat vs embedding). Mismatches return 400 preset_kind_mismatch.

How it works

Key balancing and cooldown

When you send model: "@presets/<your-slug>" to /api/v1/chat/completions or /api/v1/embeddings, the resolver picks a branch, picks a key inside that branch using your per-provider key strategy (round-robin, weighted, fallback, or random — set on the BYOK page), and dispatches the request.

If the chosen key returns 429, it's placed on a 120-second cooldown (or longer if the provider sends a Retry-After header) and the next request automatically tries the next key. If every key in every branch is exhausted, the request gets a 429 — unless you've turned on Paid fallback.

Each preset has a Paid fallback switch. When off (default), exhaustion returns 429. When on, the preset appends a synthetic "anyrouter platform" branch at the lowest priority — when every BYOK branch is unavailable, the request is served on platform credits instead. Costs land on the same cost_usd column you already see in the dashboard.

Reading the dashboard

The Capacity card at the top of every preset page shows three independent meters:

MeterWhat it measures
TokensSum of daily token budget − tokens used today across every BYOK key in the preset
RequestsSame shape for request counts
Platform $Spend vs. budget when the preset has a paid-fallback branch

A status row collapses every key into four dots:

  • ● healthy
  • ◐ cooldown (Retry-After seconds remaining)
  • ○ exhausted (≥95% of daily budget used)
  • ✕ invalid (provider rejected the key on the last probe)

A countdown row shows the next reset for each window kind, so you know when capacity comes back.

Re-probing

Limits change — providers raise free tiers, you upgrade your plan, your weekly window rolls over. Click Probe all keys on the dashboard to re-run every probe concurrently. Each probe has a 3-second budget so the page stays responsive; results land on D1 and the meters refresh on the next view.

You can also re-probe a single key via POST /api/v1/byok/probe with { "key_id": "byok_..." }.

Configure

Add your provider keys

Open Dashboard → BYOK, click Add key, choose a provider, paste the API key, give it a label, and click Validate & Add. AnyRouter validates the key against the provider and shows detected limits inline:

  • Z.ai — 5-hour token window plus per-minute request rate
  • Anthropic (Claude OAuth) — 5-hour, 7-day, and per-model windows
  • DeepSeek — USD credit balance
  • OpenRouter — USD credit balance and key-level rate limit
  • OpenAI — daily request and token counts (admin keys only)
  • Groq / Cerebras / SambaNova / NVIDIA / Mistral / Cohere / GitHub Models — per-minute request and token budget read from response headers

If a provider doesn't expose usage data, the key is saved as Unvalidated with an amber badge. You can fill in limits manually under Advanced on the same form — but you don't have to.

Build a preset

A preset is a routing tree — Preset → [Branch 1, Branch 2, …] → BYOK keys (auto-derived). Each branch is a (model, upstream) pair; the same model can appear twice via different upstreams (for example z-ai/glm-4.7-flash via your Z.ai key in one branch and the same AnyRouter model id via your OpenRouter key in another).

Open Dashboard → Presets → New. For each branch: pick a model, pick the upstream (for example Z.ai Standard API, Z.ai Coding Plan, OpenRouter, or AnyRouter platform credits), and set a weight (used in weighted mode) or priority (used in fallback mode).

Switch modes on the root:

  • Fallback — try Branch 1 first, only move to Branch 2 if every key in Branch 1 is cooled down or exhausted.
  • Weighted — distribute requests across branches by weight (e.g. 70/20/10).

Call the preset

Send model: "@presets/<your-slug>" to /api/v1/chat/completions or /api/v1/embeddings, and the resolver handles branch and key selection for you.

What's not included

A few capabilities are deliberately out of scope:

  • OAuth-subscription tiers (Claude Code, GitHub Copilot, Cursor) — they're typically tied to subscriber terms that don't allow third-party intermediation.
  • Sticky sessions — every turn re-resolves the preset; if you need turn-to-turn affinity, pin a single model in your request instead.
  • Background probing — probes run when you open the dashboard, not on a cron.

Related