Response Caching
Cache identical AI Gateway responses so repeat requests return instantly at zero upstream cost — automatic for deterministic calls, and controllable per request.
Prompt caching (see Prompt Caching) reduces the cost of a cache-read on a provider that still runs inference. Response caching goes a step further: when a request is byte-for-byte identical to one AnyRouter has already answered recently, the cached response is served directly — no upstream call at all, and no charge.
When it helps
Response caching pays off any time you send the exact same request more than once:
| Use case | Why it caches well |
|---|---|
Deterministic requests (temperature: 0) | Same input always produces the same output, so a cached response is indistinguishable from a fresh one |
| Embeddings | Embedding a given input is a pure function — the vector never changes |
| Repeated identical prompts | Health checks, smoke tests, and demo flows that resend the same payload |
| Development and debugging | Iterating on your own code around a fixed prompt without re-paying for every run |
If your requests vary even slightly (different messages, temperature, or any other field), each one is treated as distinct and caching has no effect.
Automatic default
AnyRouter automatically applies a short cache window to requests it can safely reuse:
- Deterministic chat/completions requests (
temperature: 0or unset default), embeddings, and speech requests are cached for about 5 minutes by default. - Caching is never applied to zero-data-retention (ZDR) workspaces.
- Caching is skipped whenever a request explicitly opts out (see
skipCachebelow).
You don't need to do anything to benefit from the default — it's a safety-net optimization, not a requirement.
Controlling it yourself
For more control, set the cache TTL, a custom cache key, or opt a request out entirely.
With the SDK
import { AnyRouter } from "@anyr/sdk"
const client = new AnyRouter()
// Cache this exact response for 10 minutes.
const completion = await client.chat.completions.create(
{
model: "z-ai/glm-4.7-flash",
messages: [{ role: "user", content: "What is the capital of France?" }],
},
{ cacheTtl: 600 },
)
// Share a cache entry across requests with different wording but the same
// semantic intent, by supplying your own key.
await client.chat.completions.create(
{ model: "z-ai/glm-4.7-flash", messages: [{ role: "user", content: "Hi!" }] },
{ cacheKey: "greeting-v1" },
)
// Force a fresh upstream call, bypassing any cache.
await client.chat.completions.create(
{ model: "z-ai/glm-4.7-flash", messages: [{ role: "user", content: "Latest news?" }] },
{ skipCache: true },
)
The same options are available as providerOptions.anyrouter when using @anyr/ai-sdk-provider:
import { anyrouter } from "@anyr/ai-sdk-provider"
import { generateText } from "ai"
const { text } = await generateText({
model: anyrouter("z-ai/glm-4.7-flash"),
prompt: "What is the capital of France?",
providerOptions: {
anyrouter: { cacheTtl: 600 },
},
})
With raw headers
Any HTTP client can control caching directly with request headers:
| Header | Value | Effect |
|---|---|---|
cf-aig-cache-ttl | seconds, e.g. 600 | Cache this response for the given duration |
cf-aig-cache-key | any string | Use a custom cache key instead of the request's default identity |
cf-aig-skip-cache | true | Never cache this request, even if a default policy would apply |
curl https://anyrouter.dev/api/v1/chat/completions \
-H "Authorization: Bearer $ANYROUTER_API_KEY" \
-H "Content-Type: application/json" \
-H "cf-aig-cache-ttl: 600" \
-d '{
"model": "z-ai/glm-4.7-flash",
"messages": [{ "role": "user", "content": "What is the capital of France?" }]
}'
TTL bounds
cf-aig-cache-ttl accepts a minimum of 60 seconds and a maximum of 1 month (2,678,400 seconds). Values outside that range are clamped by the gateway.