Engineering
How the gateway is built on Cloudflare Workers — routing, failover, streaming, and the edge trade-offs behind it.
2 posts

How automatic failover works in an LLM gateway
A single provider call has a single point of failure: one rate limit, one outage, one bad deploy on their end, and your request fails. Here's the actual mechanism AnyRouter uses to route around that — circuit breakers, escalating cooldowns, and a fallback chain — and why the extra hop barely registers next to model generation time.
July 12, 2026
Building an LLM gateway on Cloudflare Workers
AnyRouter runs as a set of Cloudflare Workers, not a long-running server. That constrains the design in specific ways — a 128MB memory ceiling that rules out buffering responses, a 3MiB script cap that forced a multi-Worker split, and a hard rule that every upstream call goes through Cloudflare's AI Gateway. Here's what that architecture actually looks like.
July 12, 2026