Latency Overhead
How much latency AnyRouter adds compared to calling the upstream provider directly — and when AnyRouter is actually faster.
Understand exactly what AnyRouter adds to a request, and measure it yourself. Calling an upstream API through AnyRouter adds a small, predictable amount of latency — and for some users, especially those far from the upstream's origin, AnyRouter is actually faster than calling the provider directly.
Request flow
sequenceDiagram
autonumber
participant App as Your app
participant CF as AnyRouter Worker (CF edge)
participant AIG as Cloudflare AI Gateway
participant Up as Upstream provider
App->>CF: POST /api/v1/chat/completions
Note over CF: auth + workspace lookup (D1/KV)
Note over CF: routing + cooldown check
Note over CF: body build / translate
CF->>AIG: forward request
AIG->>Up: forward request
Up-->>AIG: response / SSE stream
AIG-->>CF: response / SSE stream
CF-->>App: pipe-through (no buffering)
Note over CF: billing + audit via waitUntil (off critical path)
What AnyRouter adds
| Stage | Typical wall-clock | Notes |
|---|---|---|
| CF Worker dispatch | ~0 ms (warm) / 5–20 ms (cold) | Subsequent requests on the same isolate add nothing |
| Auth + workspace lookup | 1–10 ms | D1 read with KV cache; KV hit is sub-millisecond |
| Routing + cooldown | 1–5 ms | In-memory route table; KV cooldown check |
| Body build / translate | <1 ms | Pure JS, no I/O |
| AI Gateway hop | 10–30 ms | CF-owned; co-located with the Worker |
| Streaming pipe-through | 0 ms / byte | SSE chunks forward immediately; only the final billing frame is appended |
| Billing + audit writes | 0 ms on the critical path | All run via waitUntil after the response is flushed |
Net steady-state add: ~15–50 ms TTFB, ~10–35 ms TTFT.
When AnyRouter is faster than calling the provider directly
flowchart LR
A[Your app] -- 250ms TLS + cold connection --> B[Provider US-East]
A[Your app] -- 30ms TLS + warm pool --> C[CF edge near you]
C -- 50ms warm fan-in --> B
- You're far from the upstream origin. AnyRouter terminates TLS at the CF edge nearest you. The Worker reuses a warm connection pool to the upstream, so your packet round-trip is short and the long-haul leg is amortized across all AnyRouter traffic.
- Your own backend cold-starts. A serverless app talking to OpenAI pays its own cold-start every time it scales out. AnyRouter's Worker is almost always already warm.
- Provider rate-limits or has a transient outage. With
body.models, AnyRouter walks to the next candidate faster than your app can implement a retry loop. See Model Fallbacks. - Multi-provider redundancy. A direct integration with one provider can't fail over without code changes. AnyRouter's provider routing does it on every request.
When AnyRouter is slower
- You're already in the same region as the provider. A US-East serverless function hitting OpenAI direct will be faster than going through CF (one fewer hop).
- First request to a cold Worker isolate. Rare; subsequent requests on that isolate add nothing.
- Strict TTFB budgets on small non-streaming responses where ~20 ms matters more than reliability.
Measure it yourself
Time-to-first-byte tells you the real overhead for your location. Run the same prompt against AnyRouter and the provider directly, then compare.
# Time-to-first-byte against AnyRouter
curl -w 'ttfb=%{time_starttransfer}s total=%{time_total}s\n' \
-o /dev/null -s \
-X POST https://anyrouter.dev/api/v1/chat/completions \
-H "Authorization: Bearer <ANYROUTER_API_KEY>" \
-H "Content-Type: application/json" \
-d '{"model":"openai/gpt-5.4-mini","messages":[{"role":"user","content":"hi"}],"max_completion_tokens":10}'
# Same prompt against the provider directly (use your own key)
curl -w 'ttfb=%{time_starttransfer}s total=%{time_total}s\n' \
-o /dev/null -s \
-X POST https://api.openai.com/v1/chat/completions \
-H "Authorization: Bearer <OPENAI_API_KEY>" \
-H "Content-Type: application/json" \
-d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"hi"}],"max_tokens":10}'
Run each at least a few times to warm caches before comparing — the first request to a cold Worker isolate or a cold connection pool is not representative.