Prompt Caching Explained: Cut LLM Costs and Latency on Repeated Prompts
If every request repeats the same long system prompt, documents or tool definitions, prompt caching lets the provider reuse that work for a fraction of the price and time. How it works, structuring prompts for cache hits, Claude's cache_control, OpenAI's automatic caching, and verifying it works.
Most AI features send the same large chunk of text on every request: a long system prompt, a product manual, a set of tool definitions, the conversation so far. You pay to process that chunk every single time — and wait for it too.
Prompt caching lets the provider remember the processed form of a prompt prefix and reuse it on later requests. Cached input tokens are billed at a steep discount and processed much faster.
How it works
When a model reads a prompt, it computes internal state for every token. Prompt caching stores that state for the beginning of the prompt (the prefix). If the next request starts with exactly the same tokens, the provider loads the stored state instead of recomputing it, then processes only the new part.
Key rule: caching works on prefixes. Change anything early in the prompt, and everything after it misses the cache.
Structure prompts for cache hits
Put things in order from most stable to most variable:
- Tool definitions (rarely change)
- System prompt and instructions
- Large reference material (docs, policies, codebase context)
- Conversation history
- The new user message (changes every time)
Common cache-killers:
- A timestamp or request ID in the system prompt ("Current time: 14:03:22").
- Per-user data near the top (the user's name in the first line).
- Non-deterministic ordering — tools or documents serialised in a different order each time (e.g. iterating a map or object whose key order varies).
- Editing the system prompt during A/B tests without realising each variant has its own cache.
Move dynamic bits to the end, or into the user message.
Claude: explicit cache breakpoints
With the Claude API you mark where the cacheable prefix ends using cache_control:
await anthropic.messages.create({
model: 'claude-sonnet-5-5',
max_tokens: 1024,
system: [
{ type: 'text', text: 'You are the support assistant for Acme…' },
{
type: 'text',
text: longProductManual, // tens of thousands of tokens
cache_control: { type: 'ephemeral' }, // cache everything up to here
},
],
messages: [{ role: 'user', content: question }],
})
What to know (check Anthropic's docs for current numbers):
- The cache lives for a short time — 5 minutes by default, refreshed each time it's used; a longer 1-hour option costs more to write.
- Writing to the cache costs about 1.25× normal input (about 2× for the 1-hour option); reading from it costs about 0.1× normal input. If the prefix is reused even once or twice, you come out ahead.
- There's a minimum cacheable length — between 512 and 4,096 tokens depending on the model. Shorter prefixes silently aren't cached.
- You can set up to four breakpoints — e.g. one after tools, one after documents, one at the end of the conversation history — so partial changes still reuse earlier sections.
- If you don't need fine control, put a single top-level
cache_control: { type: 'ephemeral' }on the request and the API caches up to the last cacheable block automatically.
For multi-turn chats, putting a breakpoint on the latest message means each turn reuses the whole previous conversation.
OpenAI: automatic caching
OpenAI caches long prompt prefixes automatically — no markup — and bills cached input at a discount. The same structuring rules apply: stable content first, variable content last.
Verify it's working
Don't assume. Both APIs report cache usage in the response:
console.log(response.usage)
// Claude: cache_creation_input_tokens, cache_read_input_tokens, input_tokens
On the first request you should see cache creation; on the next identical-prefix request within the TTL, cache reads. If you only ever see creation, something in your prefix is changing. Log these numbers in production — a silent cache miss can double your bill. (Structured logging)
When caching helps most
- Chatbots with a long system prompt or knowledge base. (Add an AI chatbot)
- Multi-turn conversations — each turn re-sends the history.
- Agents with many tool definitions and long transcripts. Coding agents like Claude Code rely heavily on caching. (Keep Claude Code costs down)
- Document Q&A — many questions about the same document.
It helps little when every request is unique and short.
Caching vs other savings
- Batch processing — cheaper for work that can wait; combinable with caching.
- Smaller models — the biggest lever for simple tasks. (Claude Opus vs Sonnet vs Haiku)
- Response caching in your app — if the same question gets the same answer, store the answer and skip the model entirely. (What is caching?)
The summary
- Prompt caching reuses processing of an identical prompt prefix — cheaper and faster.
- Order prompts stable-first; keep timestamps and per-user data out of the prefix.
- Claude: mark breakpoints with
cache_control; short default lifetime, refreshed on use. - OpenAI: automatic for long prefixes.
- Check
usagefor cache reads to confirm it works.
EasySpawn runs your AI backend on its own server, where you can log token and cache usage per request and keep prompts stable across your whole app. See how it works or join the waitlist.
Related: Claude API Pricing Explained · What Are Tokens in AI? · What Is a System Prompt? · Streaming LLM Responses
Keep reading
How to Stop Bots From Running Up Your AI App's Bill
If your app calls an AI model on a user's behalf, every request costs you money — and a bot, a scraper, or one determined user can make thousands of them overnight. Rate limits, usage caps, provider spending limits, and the architecture that keeps a surprise bill from happening.
Streaming LLM Responses to the Browser: SSE, Fetch Streams, and Gotchas
Streaming makes AI features feel fast: words appear as they're generated instead of after a ten-second wait. How to stream from the Claude API on your server, forward it to the browser, read it in React, and fix the proxies and timeouts that buffer or cut off streams.