Skip to content

Rate Limits

AnyRouter's per-IP and per-key rate limits, how tiers are picked, what 429 vs 403 mean, and how to handle them in your client.

Agent loops and burst traffic can hammer an endpoint into a wall of 429s if the limits aren't clear. AnyRouter applies request-rate limits at two layers — a per-IP cap for anonymous traffic and a per-API-key cap for authenticated traffic — enforced over a sliding 60-second window using Cloudflare's native rate-limiting bindings. Paid and BYOK traffic has no daily quota (except the Go plan's per-request-window cap below); the free tier (anyrouter/free and *:free models) has a separate daily cap that resets at 00

UTC (see Free-tier daily cap below).

Overview

CallerTierLimit
Anonymous (no API key)per-IP300 req / 60s
API key, plan = go (any model, free or paid)go60 req / 60s (also capped at 3,000 req per rolling 5-hour window)
API key, free-tier model (*:free), plan = free, payg, pro, pro_plus or maxfree-model240 req / 60s
API key, plan = free or payg, paid modelpaid-model600 req / 60s
API key, plan = pro, paid modelpro1200 req / 60s
API key, plan = pro_plus or max, paid modelpower1800 req / 60s
API key, plan = enterprise (any model)enterprise2400 req / 60s

Go is the one exception to the free-model row above. Go's 60 req/min ceiling applies to every model — free-tier included. A Go-plan key calling anyrouter/free is limited at 60/min, not 240/min. On entitled plans (pay-as-you-go with credits, Pro, Pro+, Max, Enterprise), free-tier models use the 240/min tier.

The X-RateLimit-Tier value for Pro paid-model traffic is pro; for Pro+ and Max it is power — legacy identifiers kept for backward compatibility. Account plan names are Free, Go, Pro, Pro+, and Max. See Pricing & Plans.

Why the split: free models are subsidised, so we cap that traffic tightly to prevent abuse. Paid models are billed per token — your spend naturally regulates volume, so we lift the per-minute cap on Pro and above. Go keeps a tighter per-minute and 5-hour window cap since it's the entry paid tier. Enterprise plans get the same headroom regardless of which model class they hit.

Your account plan (free / go / pro / pro_plus / max / enterprise) determines your rate tier for paid-model traffic. A Free user calling a paid model uses the paid-model tier at 600/min; Go is capped at 60/min (plus a 3,000 req/5h window), Pro at 1200/min, and Pro+ / Max at 1800/min. Free-tier models run at the shared 240/min free-model tier on Free, pay-as-you-go, Pro, Pro+, and Max — except Go (60/min on every model) and Enterprise (2400/min on every model) — with the per-account daily cap below on top (10/day on Free, 1000/day on paid plans).

How it works

Free-tier daily cap

The anyrouter/free model and any model suffixed with :free share a daily cap per account on top of the per-minute rate limit above: 10 requests/day on Free, 1000 requests/day on Go, Pro, Pro+, or Max. The counter resets at 00

UTC every day. Paid models do not consume this cap. See Free Tier.

This cap only applies to platform-managed free-tier traffic. Paid models, requests served through a BYOK key (even for :free-suffixed models), and requests that consume credits are not subject to the daily cap.

When you exceed the daily cap, the API returns:

HTTP/1.1 429 Too Many Requests
Retry-After: <seconds until 00:00 UTC>
Content-Type: application/json
{
  "error": {
    "code": "free_tier_daily_limit_exceeded",
    "message": "Free-tier daily limit reached. Resets at 00:00 UTC.",
    "type": "rate_limit_error"
  }
}

Read the Retry-After header to know how many seconds remain until midnight UTC. Retrying before then returns the same 429 — the cap does not refresh mid-day. To get past it, wait for the daily reset, switch to a paid model, or add credits.

Response headers

Every API response — including GET /api/v1/health, GET /api/v1/models, GET /api/v1/credits, error bodies, and successful completions — carries:

HeaderMeaning
X-RateLimit-LimitPer-window request cap for the tier you hit
X-RateLimit-WindowWindow length in seconds (always 60)
X-RateLimit-TierOne of ip, free-model, paid-model, go, pro, power, enterprise, or starter-window (Go's 5-hour window cap — see below)
RateLimit-LimitIETF draft header — same number as X-RateLimit-Limit
RateLimit-PolicyIETF draft policy, e.g. 300;w=60 (cap ; window seconds)

The per-minute tier does not emit X-RateLimit-Remaining / X-RateLimit-Reset — the counter is not exposed by the edge limiter. The two D1-counted limiters DO emit them: free-tier daily responses carry X-RateLimit-Remaining, X-RateLimit-Reset and the matching RateLimit-Remaining / RateLimit-Reset (window 86400), and Go's 5-hour window carries the same pair (window 18000, tier starter-window).

When you hit the cap, the response is HTTP 429 with a Retry-After: 60 header and a body of {"error": {"code": "rate_limit_exceeded", "message": "Rate limit exceeded for your API key — …"}}.

429 vs 403 — don't confuse them

AnyRouter returns 429 for three distinct rate-limit scenarios, each with its own error.code: rate_limit_exceeded (per-key per-minute cap), ip_rate_limit_exceeded (anonymous IP cap), and free_tier_daily_limit_exceeded (free-tier daily cap). If you see a different status:

Statuserror.codeSourceWhat it means
429rate_limit_exceededAnyRouterYou hit the per-key limit. Back off Retry-After seconds.
429ip_rate_limit_exceededAnyRouterAnonymous IP hit 300/min. Sign in or send an API key.
429free_tier_daily_limit_exceededAnyRouterFree-tier daily cap reached (10 req/day on Free, 1000 req/day on Go/Pro/Pro+/Max); resets at 00
UTC. Add credits or subscribe.
403upstream_403Upstream providerThe model provider denied the call — geo block, policy violation, provider-side abuse heuristic. Not retryable via simple backoff.
401upstream_401Upstream providerBYOK credential rejected upstream. Check your key.
401invalid_api_keyAnyRouterYour AnyRouter API key is wrong or revoked.
402insufficient_balanceAnyRouterYour credit balance is empty. Top up at Credits.
502upstream_5xxUpstream providerProvider had a transient failure. Retry with jitter.

error.metadata.upstream_backend identifies the specific provider when the status is upstream_*.

Configure

Handling 429 in your client

Check error.code before deciding how to retry:

  • rate_limit_exceeded / ip_rate_limit_exceeded — Retry-After is 60 seconds. Back off and retry.
  • free_tier_daily_limit_exceeded — Retry-After is up to 86,400 seconds (seconds until midnight UTC). Do not loop-retry; surface the message to your user or switch to a paid model.
async function callWithBackoff(body: object): Promise<Response> {
  for (let attempt = 0; attempt < 5; attempt++) {
    const res = await fetch("https://anyrouter.dev/api/v1/chat/completions", {
      method: "POST",
      headers: {
        "Authorization": `Bearer ${ANYROUTER_KEY}`,
        "Content-Type": "application/json",
      },
      body: JSON.stringify(body),
    })
    if (res.status !== 429) return res

    const err = await res.clone().json().catch(() => null)
    // Daily cap: Retry-After can be hours — surface to caller instead of sleeping
    if (err?.error?.code === "free_tier_daily_limit_exceeded") {
      throw new Error(err.error.message)
    }

    const wait = Number(res.headers.get("Retry-After") ?? 60)
    const jitter = Math.random() * 1000
    await new Promise((r) => setTimeout(r, wait * 1000 + jitter))
  }
  throw new Error("Rate-limit retries exhausted")
}

Do not treat a 429 as a permanent credential failure. Many agent loops mark keys as "exhausted" on 429 and stop using them — that's a client bug, not an AnyRouter behaviour. For per-minute 429s, the same key works again after the 60-second window.

Agent loops count every call

A single user task that triggers an agent loop can issue many HTTP requests:

  • 1 main chat/completions call
  • 1 more call per tool round-trip (one for each model turn after a tool result)
  • Optional auxiliary calls — title generation, conversation compaction, summarisation

Each is metered separately. A 5-tool task with title-gen + compaction is ~8 requests. At the 600/min paid-model tier you can sustain ~75 such tasks per minute; at the 240/min free-model tier, ~30.

If your agent's natural request rate exceeds the tier:

  1. Call paid models (auto-promotes you to 600/min).
  2. Upgrade to enterprise for 2400/min.
  3. Add jitter and parallelism caps in your loop.

Need more headroom?

Open a request at /contact describing your workload (sustained req/min, burst pattern). Enterprise tiers are negotiated.

Related