Rate Limits
AnyRouter's per-IP and per-key rate limits, how tiers are picked, what 429 vs 403 mean, and how to handle them in your client.
Agent loops and burst traffic can hammer an endpoint into a wall of 429s if the limits aren't clear. AnyRouter applies request-rate limits at two layers — a per-IP cap for anonymous traffic and a per-API-key cap for authenticated traffic — enforced over a sliding 60-second window using Cloudflare's native rate-limiting bindings. Paid and BYOK traffic has no daily quota (except the Go plan's per-request-window cap below); the free tier (anyrouter/free and *:free models) has a separate daily cap that resets at 00
Overview
| Caller | Tier | Limit |
|---|---|---|
| Anonymous (no API key) | per-IP | 300 req / 60s |
API key, plan = go (any model, free or paid) | go | 60 req / 60s (also capped at 3,000 req per rolling 5-hour window) |
API key, free-tier model (*:free), plan = free, payg, pro, pro_plus or max | free-model | 240 req / 60s |
API key, plan = free or payg, paid model | paid-model | 600 req / 60s |
API key, plan = pro, paid model | pro | 1200 req / 60s |
API key, plan = pro_plus or max, paid model | power | 1800 req / 60s |
API key, plan = enterprise (any model) | enterprise | 2400 req / 60s |
Go is the one exception to the free-model row above. Go's 60 req/min ceiling applies to every model — free-tier included. A Go-plan key calling anyrouter/free is limited at 60/min, not 240/min. On entitled plans (pay-as-you-go with credits, Pro, Pro+, Max, Enterprise), free-tier models use the 240/min tier.
The X-RateLimit-Tier value for Pro paid-model traffic is pro; for Pro+ and Max it is power — legacy identifiers kept for backward compatibility. Account plan names are Free, Go, Pro, Pro+, and Max. See Pricing & Plans.
Why the split: free models are subsidised, so we cap that traffic tightly to prevent abuse. Paid models are billed per token — your spend naturally regulates volume, so we lift the per-minute cap on Pro and above. Go keeps a tighter per-minute and 5-hour window cap since it's the entry paid tier. Enterprise plans get the same headroom regardless of which model class they hit.
Your account plan (free / go / pro / pro_plus / max / enterprise) determines your rate tier for paid-model traffic. A Free user calling a paid model uses the paid-model tier at 600/min; Go is capped at 60/min (plus a 3,000 req/5h window), Pro at 1200/min, and Pro+ / Max at 1800/min. Free-tier models run at the shared 240/min free-model tier on Free, pay-as-you-go, Pro, Pro+, and Max — except Go (60/min on every model) and Enterprise (2400/min on every model) — with the per-account daily cap below on top (10/day on Free, 1000/day on paid plans).
How it works
Free-tier daily cap
The anyrouter/free model and any model suffixed with :free share a daily cap per account on top of the per-minute rate limit above: 10 requests/day on Free, 1000 requests/day on Go, Pro, Pro+, or Max. The counter resets at 00
This cap only applies to platform-managed free-tier traffic. Paid models, requests served through a BYOK key (even for :free-suffixed models), and requests that consume credits are not subject to the daily cap.
When you exceed the daily cap, the API returns:
HTTP/1.1 429 Too Many Requests
Retry-After: <seconds until 00:00 UTC>
Content-Type: application/json
{
"error": {
"code": "free_tier_daily_limit_exceeded",
"message": "Free-tier daily limit reached. Resets at 00:00 UTC.",
"type": "rate_limit_error"
}
}
Read the Retry-After header to know how many seconds remain until midnight UTC. Retrying before then returns the same 429 — the cap does not refresh mid-day. To get past it, wait for the daily reset, switch to a paid model, or add credits.
Response headers
Every API response — including GET /api/v1/health, GET /api/v1/models,
GET /api/v1/credits, error bodies, and successful completions — carries:
| Header | Meaning |
|---|---|
X-RateLimit-Limit | Per-window request cap for the tier you hit |
X-RateLimit-Window | Window length in seconds (always 60) |
X-RateLimit-Tier | One of ip, free-model, paid-model, go, pro, power, enterprise, or starter-window (Go's 5-hour window cap — see below) |
RateLimit-Limit | IETF draft header — same number as X-RateLimit-Limit |
RateLimit-Policy | IETF draft policy, e.g. 300;w=60 (cap ; window seconds) |
The per-minute tier does not emit X-RateLimit-Remaining / X-RateLimit-Reset —
the counter is not exposed by the edge limiter. The two D1-counted limiters DO
emit them: free-tier daily responses carry X-RateLimit-Remaining,
X-RateLimit-Reset and the matching RateLimit-Remaining / RateLimit-Reset
(window 86400), and Go's 5-hour window carries the same pair (window
18000, tier starter-window).
When you hit the cap, the response is HTTP 429 with a Retry-After: 60 header and a body of {"error": {"code": "rate_limit_exceeded", "message": "Rate limit exceeded for your API key — …"}}.
429 vs 403 — don't confuse them
AnyRouter returns 429 for three distinct rate-limit scenarios, each with its own error.code: rate_limit_exceeded (per-key per-minute cap), ip_rate_limit_exceeded (anonymous IP cap), and free_tier_daily_limit_exceeded (free-tier daily cap). If you see a different status:
| Status | error.code | Source | What it means |
|---|---|---|---|
429 | rate_limit_exceeded | AnyRouter | You hit the per-key limit. Back off Retry-After seconds. |
429 | ip_rate_limit_exceeded | AnyRouter | Anonymous IP hit 300/min. Sign in or send an API key. |
429 | free_tier_daily_limit_exceeded | AnyRouter | Free-tier daily cap reached (10 req/day on Free, 1000 req/day on Go/Pro/Pro+/Max); resets at 00 UTC. Add credits or subscribe. |
403 | upstream_403 | Upstream provider | The model provider denied the call — geo block, policy violation, provider-side abuse heuristic. Not retryable via simple backoff. |
401 | upstream_401 | Upstream provider | BYOK credential rejected upstream. Check your key. |
401 | invalid_api_key | AnyRouter | Your AnyRouter API key is wrong or revoked. |
402 | insufficient_balance | AnyRouter | Your credit balance is empty. Top up at Credits. |
502 | upstream_5xx | Upstream provider | Provider had a transient failure. Retry with jitter. |
error.metadata.upstream_backend identifies the specific provider when the status is upstream_*.
Configure
Handling 429 in your client
Check error.code before deciding how to retry:
rate_limit_exceeded/ip_rate_limit_exceeded—Retry-Afteris 60 seconds. Back off and retry.free_tier_daily_limit_exceeded—Retry-Afteris up to 86,400 seconds (seconds until midnight UTC). Do not loop-retry; surface the message to your user or switch to a paid model.
async function callWithBackoff(body: object): Promise<Response> {
for (let attempt = 0; attempt < 5; attempt++) {
const res = await fetch("https://anyrouter.dev/api/v1/chat/completions", {
method: "POST",
headers: {
"Authorization": `Bearer ${ANYROUTER_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify(body),
})
if (res.status !== 429) return res
const err = await res.clone().json().catch(() => null)
// Daily cap: Retry-After can be hours — surface to caller instead of sleeping
if (err?.error?.code === "free_tier_daily_limit_exceeded") {
throw new Error(err.error.message)
}
const wait = Number(res.headers.get("Retry-After") ?? 60)
const jitter = Math.random() * 1000
await new Promise((r) => setTimeout(r, wait * 1000 + jitter))
}
throw new Error("Rate-limit retries exhausted")
}
Do not treat a 429 as a permanent credential failure. Many agent loops mark keys as "exhausted" on 429 and stop using them — that's a client bug, not an AnyRouter behaviour. For per-minute 429s, the same key works again after the 60-second window.
Agent loops count every call
A single user task that triggers an agent loop can issue many HTTP requests:
- 1 main
chat/completionscall - 1 more call per tool round-trip (one for each model turn after a tool result)
- Optional auxiliary calls — title generation, conversation compaction, summarisation
Each is metered separately. A 5-tool task with title-gen + compaction is ~8 requests. At the 600/min paid-model tier you can sustain ~75 such tasks per minute; at the 240/min free-model tier, ~30.
If your agent's natural request rate exceeds the tier:
- Call paid models (auto-promotes you to 600/min).
- Upgrade to enterprise for 2400/min.
- Add jitter and parallelism caps in your loop.
Need more headroom?
Open a request at /contact describing your workload (sustained req/min, burst pattern). Enterprise tiers are negotiated.