Prompt Caching
Cache long system prompts and reusable context across requests to cut cost and latency on conversational, RAG, and agent workloads.
Resending the same multi-thousand-token system prompt on every turn means paying full input price for tokens the model has already seen. AnyRouter supports prompt caching on every upstream that exposes it (Anthropic, OpenAI, DeepSeek, Gemini): cache a long prefix once and reuse it across thousands of requests, with cached tokens billed at a fraction of the normal input price.
Overview
Prompt caching pays off any time the same prefix is sent many times. Concrete wins:
| Use case | What to cache |
|---|---|
| System prompts | A multi-thousand-token persona or tool spec repeated on every turn |
| RAG | Retrieved documents that are the same for a batch of questions |
| Agent scaffolds | Long tool catalogs shared across every step of a chain |
| Few-shot examples | Static demonstrations pinned to the front of the prompt |
Rule of thumb: if a chunk is repeated across two or more requests within five minutes, cache it. The break-even vs. uncached input is typically 1.5–2 uses.
How it works
Anthropic caching
Anthropic supports breakpoint-style caching. Mark blocks with cache_control:
const response = await client.messages.create({
model: "anthropic/claude-sonnet-4.6",
max_tokens: 1024,
system: [
{
type: "text",
text: LONG_SYSTEM_PROMPT,
cache_control: { type: "ephemeral" },
},
],
messages: [{ role: "user", content: "What's the policy on refunds?" }],
})
Cache reads are billed at 10% of input price; cache writes are billed at 125%. The cache TTL is 5 minutes.
OpenAI automatic caching
OpenAI caches prompts ≥1024 tokens automatically — no request changes needed. The first request warms the cache; subsequent requests within the cache window pay reduced input pricing on cached tokens.
Pricing
Every model's detail page shows both standard and cache-read pricing. Cached tokens are reflected in the usage field of the response:
{
"usage": {
"prompt_tokens": 5840,
"completion_tokens": 120,
"total_tokens": 5960,
"prompt_tokens_details": {
"cached_tokens": 5800
}
}
}
AnyRouter bills cached tokens at the provider's cache-read rate, pass-through — no extra AnyRouter markup on cache reads.