Skip to content

Latency Overhead

How much latency AnyRouter adds compared to calling the upstream provider directly — and when AnyRouter is actually faster.

Understand exactly what AnyRouter adds to a request, and measure it yourself. Calling an upstream API through AnyRouter adds a small, predictable amount of latency — and for some users, especially those far from the upstream's origin, AnyRouter is actually faster than calling the provider directly.

Request flow

sequenceDiagram
  autonumber
  participant App as Your app
  participant CF as AnyRouter Worker (CF edge)
  participant AIG as Cloudflare AI Gateway
  participant Up as Upstream provider

  App->>CF: POST /api/v1/chat/completions
  Note over CF: auth + workspace lookup (D1/KV)
  Note over CF: routing + cooldown check
  Note over CF: body build / translate
  CF->>AIG: forward request
  AIG->>Up: forward request
  Up-->>AIG: response / SSE stream
  AIG-->>CF: response / SSE stream
  CF-->>App: pipe-through (no buffering)
  Note over CF: billing + audit via waitUntil (off critical path)

What AnyRouter adds

StageTypical wall-clockNotes
CF Worker dispatch~0 ms (warm) / 5–20 ms (cold)Subsequent requests on the same isolate add nothing
Auth + workspace lookup1–10 msD1 read with KV cache; KV hit is sub-millisecond
Routing + cooldown1–5 msIn-memory route table; KV cooldown check
Body build / translate<1 msPure JS, no I/O
AI Gateway hop10–30 msCF-owned; co-located with the Worker
Streaming pipe-through0 ms / byteSSE chunks forward immediately; only the final billing frame is appended
Billing + audit writes0 ms on the critical pathAll run via waitUntil after the response is flushed

Net steady-state add: ~15–50 ms TTFB, ~10–35 ms TTFT.

When AnyRouter is faster than calling the provider directly

flowchart LR
  A[Your app] -- 250ms TLS + cold connection --> B[Provider US-East]
  A[Your app] -- 30ms TLS + warm pool --> C[CF edge near you]
  C -- 50ms warm fan-in --> B
  • You're far from the upstream origin. AnyRouter terminates TLS at the CF edge nearest you. The Worker reuses a warm connection pool to the upstream, so your packet round-trip is short and the long-haul leg is amortized across all AnyRouter traffic.
  • Your own backend cold-starts. A serverless app talking to OpenAI pays its own cold-start every time it scales out. AnyRouter's Worker is almost always already warm.
  • Provider rate-limits or has a transient outage. With body.models, AnyRouter walks to the next candidate faster than your app can implement a retry loop. See Model Fallbacks.
  • Multi-provider redundancy. A direct integration with one provider can't fail over without code changes. AnyRouter's provider routing does it on every request.

When AnyRouter is slower

  • You're already in the same region as the provider. A US-East serverless function hitting OpenAI direct will be faster than going through CF (one fewer hop).
  • First request to a cold Worker isolate. Rare; subsequent requests on that isolate add nothing.
  • Strict TTFB budgets on small non-streaming responses where ~20 ms matters more than reliability.

Measure it yourself

Time-to-first-byte tells you the real overhead for your location. Run the same prompt against AnyRouter and the provider directly, then compare.

# Time-to-first-byte against AnyRouter
curl -w 'ttfb=%{time_starttransfer}s total=%{time_total}s\n' \
  -o /dev/null -s \
  -X POST https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer <ANYROUTER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{"model":"openai/gpt-5.4-mini","messages":[{"role":"user","content":"hi"}],"max_completion_tokens":10}'
# Same prompt against the provider directly (use your own key)
curl -w 'ttfb=%{time_starttransfer}s total=%{time_total}s\n' \
  -o /dev/null -s \
  -X POST https://api.openai.com/v1/chat/completions \
  -H "Authorization: Bearer <OPENAI_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"hi"}],"max_tokens":10}'

Run each at least a few times to warm caches before comparing — the first request to a cold Worker isolate or a cold connection pool is not representative.

Related