Health API
GET /api/v1/health — live snapshot with status, checked_at, and per-component health.
Return a public snapshot of AnyRouter's runtime health: overall status, when it was checked, and a components breakdown. No API key required. Wire it into status pages or probes.
/api/v1/healthThis endpoint is unauthenticated. A default curl User-Agent may receive 403 HTML from zone bot protection on the apex — use https://api.anyrouter.dev/api/v1/health, or send a descriptive User-Agent such as -A "my-status-probe/1.0".
Request
No parameters and no auth header. Recommended polling interval: ≥ 60 s (see Caching).
GET /api/v1/health HTTP/1.1
Prefer the dedicated API host if a client (including default curl) receives HTML instead of JSON:
curl https://api.anyrouter.dev/api/v1/health
https://anyrouter.dev/api/v1/health is the same Worker.
Response
The handler always returns HTTP 200 on success. Read status in the JSON body — do not treat the HTTP code as the health signal.
Each component has its own status. They are independent:
components.apiis last-hour generation error rate (can bedownduring a traffic spike). BYOK requests (generations.is_byok = 1) are excluded entirely — a user's own upstream key is not a platform signal. A failed request only counts as an error when its final attempt was a real platform outage (404/410, any 5xx, orfailure_classhealth/capability) — a key-permission failure (401/403,credential) never counts.components.providerslists routing backend ids that served traffic in the last hour (cloudflare, notopenrouter-byok) — not catalog model owners (openai,z-ai), and not-byok/-poolbackends (a user's or donor's own key is their problem, never a platform incident). Idle backends are omitted. Last-hour generations whoseprovider_idis not a public backend appear asunattributedso this list andcomponents.api.data.requests_1hshare one window. Quarantine, budget-exhausted, and smoke-exclusion overlay the traffic-derived status so a skipped backend cannot readhealthy. Error counting follows the same real-outage-only rule ascomponents.api.components.routing.summary.healthyis how many upstreams are currently routable — it subtracts quarantined and budget-exhausted backends, so it will not readN/Nwhile those suppressions are active.components.smoke_tests.last_runis the daily catalog smoke (00 UTC). A run from earlier today is expected;data.staleis true only whenlast_runis older than two days (2× the daily cadence).eventsis a chronological list of upstream status-change events (not every probe tick), newest first, capped at 50.-byok/-poolupstreams never appear here.components.postgresandcomponents.clickhouseare optional archives. They report probe status from the homelab latency check. They never move overallstatus— inference, API keys, and billing stay on D1, Analytics Engine, and R2.
{
"status": "healthy",
"description": "All systems operational.",
"version": "1.4.0",
"checked_at": "2026-08-18T12:00:00.000Z",
"components": {
"api": {
"status": "healthy",
"description": "Processing requests normally. 12847 requests in the last hour with 2.1% error rate.",
"note": "Derived from the last 1 hour of traffic.",
"data": {
"requests_1h": 12847,
"error_rate": 0.021,
"latency_ms": { "avg": 340 },
"cache_hit_rate": 0.35
}
},
"providers": {
"status": "healthy",
"description": "All 1 providers that served traffic are operational.",
"note": "Each row is a routing backend id (`generations.provider_id`), not a catalog model owner. Status is the last-hour error rate, overlaid with quarantine / budget-exhausted / smoke-exclusion.",
"data": {
"providers": [
{
"name": "cloudflare",
"status": "healthy",
"description": "Operating normally (5421 requests, 1.2% errors).",
"latency_ms": 120,
"requests_1h": 5421,
"error_rate": 0.012
}
]
}
},
"smoke_tests": {
"status": "healthy",
"description": "All 80 tested models passed.",
"note": "The daily smoke-test workflow (00:00 UTC) sends a real one-token request through chat/completions, responses, and messages.",
"data": {
"total": 80,
"passed": 80,
"failed": 0,
"skipped": 0,
"last_run": "2026-08-18T00:00:00.000Z",
"stale": false,
"cadence": "0 0 * * *"
}
},
"workflows": {
"status": "healthy",
"description": "All 8 scheduled workflows are running on cadence.",
"note": "Each scheduled workflow that has recorded a run is checked against its last run. Jobs that have never run are omitted.",
"data": {
"healthy": 8,
"stale": 0,
"failing": 0,
"never_ran": 0,
"workflows": [
{
"kind": "gateway-log-sync",
"label": "Gateway log sync",
"state": "healthy",
"last_status": "success",
"last_run_at": "2026-08-18T11:45:00.000Z",
"age_hours": 0.1,
"detail": "Last run success 0.1h ago."
}
]
}
},
"routing": {
"status": "healthy",
"description": "All 84 routing upstreams are healthy.",
"note": "Live routing health from the smoke-cooldown store, plus active quarantine and monthly-budget suppression.",
"data": {
"summary": {
"total": 84,
"healthy": 84,
"deprioritized": 0,
"excluded": 0,
"quarantined": 0,
"budget_exhausted": 0
},
"degraded": []
}
}
},
"events": [
{
"at": "2026-08-18T11:10:00.000Z",
"upstream_id": "cerebras",
"from": "healthy",
"to": "degraded",
"message": "HTTP 503: Service Unavailable"
}
]
}
| Field | Type | Description |
|---|---|---|
status | "healthy" | "degraded" | "partial" | "down" | "no_signal" | Overall state. Same enum is used on every component. partial means some components are degraded while routing still works. no_signal means there is not enough recent traffic to assess. |
description | string | Human-readable summary of the overall state. |
version | string | Release tag of the currently-deployed Worker. |
checked_at | string | ISO 8601 timestamp of this snapshot. |
components | object | Breakdown for api, providers, smoke_tests, workflows, routing, postgres, and clickhouse. smoke_tests, workflows, and routing may be null. Postgres and ClickHouse are optional archives and never move overall status. |
events | array | Chronological upstream status-change events, newest first (capped at 50). Never every probe tick — only rows where the status changed from the previous check. |
Every component object has status (same enum), description, note (how the check is derived), and data.
| Nested field | Type | Description |
|---|---|---|
components.api.data.requests_1h | integer | Requests in the last hour. |
components.api.data.error_rate | number | Fraction of those requests that errored (0–1). |
components.api.data.latency_ms.avg | number | Average latency in milliseconds. |
components.api.data.cache_hit_rate | number | Fraction of requests that hit the prompt cache (0–1). |
components.providers.data.providers[] | array | Per-backend rows (name is a routing backend id such as cloudflare, not a catalog owner). Includes status, description, optional latency_ms / error_rate, and requests_1h. Only backends with last-hour traffic appear. Not a top-level providers array. |
components.smoke_tests.data | object | total, passed, failed, skipped, last_run (ISO 8601 or null), stale (true only after 2× the daily cadence), and cadence (0 0 * * *). |
components.workflows.data | object | Counts healthy, stale, failing, never_ran, plus workflows[] (kind, label, state, last_status, last_run_at, age_hours, detail). Public health omits jobs that have never run, so never_ran is always 0 and workflows[] has no never_ran rows. Jobs whose schedule is intentionally paused are omitted for the same reason — a paused job is not failing, and it has not proven it runs on cadence either. |
components.routing.data.summary | object | total, healthy, deprioritized, excluded, quarantined, budget_exhausted. healthy subtracts quarantined / budget-exhausted backends that still look smoke-healthy. |
components.routing.data.degraded[] | array | Degraded upstreams with id (routing backend id, e.g. google-agent-platform), name (display label — not unique), logo, and state (deprioritized, excluded, quarantined, or budget_exhausted). A backend degraded in more than one way (for example both quarantined and out of budget) appears once per state. |
events[] | array | Each entry has at (ISO 8601), upstream_id, from/to (healthy, degraded, down, or unknown), and an optional message from the probe. |
A binary probe can treat any top-level status other than healthy as not fully operational. Pair provider rows with /api/v1/providers for logos and descriptions.
Responses also include X-RateLimit-Limit, X-RateLimit-Window, and X-RateLimit-Tier (anonymous callers use the IP tier). A 429 honors Retry-After. See Rate Limits.
Examples
curl https://api.anyrouter.dev/api/v1/health
import httpx
health = httpx.get("https://api.anyrouter.dev/api/v1/health").json()
if health["status"] != "healthy":
print(health["description"], health["components"]["providers"]["status"])
const res = await fetch("https://api.anyrouter.dev/api/v1/health")
const health = await res.json()
if (health.status !== "healthy") {
console.warn(health.description, health.components)
}
Status codes
| HTTP | Body status | Meaning |
|---|---|---|
| 200 | healthy | All scored components are healthy. |
| 200 | partial / degraded / no_signal / down | One or more components are unhealthy or lack traffic. Inference may still succeed via fallback. |
| 429 | — | Anonymous IP cap (300 req / 60s). Honor Retry-After. |
This endpoint does not use HTTP 503 for a down component — read status and components.*.status.
Caching
Anonymous responses are cacheable. The Worker stamps:
Cache-Control: public, max-age=60, s-maxage=300, stale-while-revalidate=600
Browsers may reuse the body for 60 seconds; the shared edge may keep it for 5 minutes and serve a stale copy for up to 10 minutes while one request refreshes it. Poll at ≥ 60 s.
Related
- Providers — catalog of upstream backends
- Models — the catalog
- Rate Limits —
X-RateLimit-*headers - API Overview — every endpoint at a glance