Skip to content

Response Caching

Cache identical AI Gateway responses so repeat requests return instantly at zero upstream cost — automatic for deterministic calls, and controllable per request.

Prompt caching (see Prompt Caching) reduces the cost of a cache-read on a provider that still runs inference. Response caching goes a step further: when a request is byte-for-byte identical to one AnyRouter has already answered recently, the cached response is served directly — no upstream call at all, and no charge.

When it helps

Response caching pays off any time you send the exact same request more than once:

Use caseWhy it caches well
Deterministic requests (temperature: 0)Same input always produces the same output, so a cached response is indistinguishable from a fresh one
EmbeddingsEmbedding a given input is a pure function — the vector never changes
Repeated identical promptsHealth checks, smoke tests, and demo flows that resend the same payload
Development and debuggingIterating on your own code around a fixed prompt without re-paying for every run

If your requests vary even slightly (different messages, temperature, or any other field), each one is treated as distinct and caching has no effect.

Automatic default

AnyRouter automatically applies a short cache window to requests it can safely reuse:

  • Deterministic chat/completions requests (temperature: 0 or unset default), embeddings, and speech requests are cached for about 5 minutes by default.
  • Caching is never applied to zero-data-retention (ZDR) workspaces.
  • Caching is skipped whenever a request explicitly opts out (see skipCache below).

You don't need to do anything to benefit from the default — it's a safety-net optimization, not a requirement.

Controlling it yourself

For more control, set the cache TTL, a custom cache key, or opt a request out entirely.

With the SDK

import { AnyRouter } from "@anyr/sdk"

const client = new AnyRouter()

// Cache this exact response for 10 minutes.
const completion = await client.chat.completions.create(
  {
    model: "z-ai/glm-4.7-flash",
    messages: [{ role: "user", content: "What is the capital of France?" }],
  },
  { cacheTtl: 600 },
)

// Share a cache entry across requests with different wording but the same
// semantic intent, by supplying your own key.
await client.chat.completions.create(
  { model: "z-ai/glm-4.7-flash", messages: [{ role: "user", content: "Hi!" }] },
  { cacheKey: "greeting-v1" },
)

// Force a fresh upstream call, bypassing any cache.
await client.chat.completions.create(
  { model: "z-ai/glm-4.7-flash", messages: [{ role: "user", content: "Latest news?" }] },
  { skipCache: true },
)

The same options are available as providerOptions.anyrouter when using @anyr/ai-sdk-provider:

import { anyrouter } from "@anyr/ai-sdk-provider"
import { generateText } from "ai"

const { text } = await generateText({
  model: anyrouter("z-ai/glm-4.7-flash"),
  prompt: "What is the capital of France?",
  providerOptions: {
    anyrouter: { cacheTtl: 600 },
  },
})

With raw headers

Any HTTP client can control caching directly with request headers:

HeaderValueEffect
cf-aig-cache-ttlseconds, e.g. 600Cache this response for the given duration
cf-aig-cache-keyany stringUse a custom cache key instead of the request's default identity
cf-aig-skip-cachetrueNever cache this request, even if a default policy would apply
curl https://anyrouter.dev/api/v1/chat/completions \
  -H "Authorization: Bearer $ANYROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -H "cf-aig-cache-ttl: 600" \
  -d '{
    "model": "z-ai/glm-4.7-flash",
    "messages": [{ "role": "user", "content": "What is the capital of France?" }]
  }'

TTL bounds

cf-aig-cache-ttl accepts a minimum of 60 seconds and a maximum of 1 month (2,678,400 seconds). Values outside that range are clamped by the gateway.

Related