Imagine consulting a lawyer every week on the same case. At each appointment, they re-read the entire file from the start, and bill you for those hours of reading. Absurd, but that's exactly what happens when you send the same long context to an AI with every request: you pay for the full re-read, over and over. Prompt caching is telling the lawyer to keep the file open on their desk.
It's probably the most cost-effective and least-used lever in enterprise AI. With the price increases announced for September on several models, including Claude Sonnet 5, now is a good time to understand how it works.
The problem it solves
An AI has no memory between calls. With each request, you have to resend everything it needs: the system prompt, instructions, reference documents, conversation history. That's what's called the context, explained in our dedicated article.
Yet this context is often largely identical from one call to the next. An agent working on a codebase resends the same codebase at every step. A business assistant resends the same three-thousand-word system prompt with every question. A document chatbot resends the same reference documents constantly.
And you pay for those input tokens at full price, every time. On an agent chaining fifty steps, you've billed fifty times for reading the same document.
The mechanism
Prompt caching lets you mark a portion of your context as stable. The provider then keeps a processed representation of that portion for a limited time. On subsequent calls, if that portion is identical, it isn't fully reprocessed: it's read from the cache, at a much lower rate.
The savings are substantial. On Claude models, for example, cached reads are billed at a fraction of the normal input rate, with discounts of up to 90% on repeated tokens. Some providers go further: at Moonshot, re-reading content already in cache is even free.
Watch out for a subtlety. Putting something in cache has an initial cost above the normal rate: expect around 1.25 times the input rate for a short-lived cache, and up to 2 times for a one-hour cache. The math is then simple: caching pays off if you re-read that content several times before it expires. Over two calls, it's not worth it. Over fifty, it's massive. For content you use only once, it's a dead loss.
The three rules to make it work
1. Put the stable stuff at the start, the variable at the end. This is the most important and most misunderstood rule. The cache works on an identical prefix: if you change a single character at the beginning of your context, everything after it becomes invalid. So structure your requests by placing the system prompt and reference documents at the top, and the user's question right at the end.
2. Watch the lifetime. A cache expires, usually after a few minutes if not reused. If your calls are twenty minutes apart, a five-minute cache will never help you, and you'll have paid the write surcharge for nothing.
3. Verify it actually works. API responses show how many tokens were read from cache and how many were written. That's the only proof your setup is working. Many teams think they've enabled caching when a detail invalidates the prefix on every call, and they pay the write surcharge without ever benefiting from the read.
Why it's especially relevant now
Three developments make this lever more important than before.
Price increases first, with several promotional periods ending at the end of summer. A well-built cache can offset a significant chunk of a 50% increase.
Autonomous agents next. As we explained in our article on agents, these systems work in a loop: observe, plan, act, verify. Each turn resends most of the previous context. That's the ideal use case for caching, and also the one where forgetting it costs the most.
Giant contexts finally. When models accept a million tokens, the temptation is to dump entire documents in. Without caching, that's financially untenable.
What to remember
Prompt caching is a typical example of invisible optimisation: it changes nothing about the quality of your results, only the size of your bill. That's also why it's often overlooked, since nothing breaks when you forget it.
The useful reflex is simple: every time you're about to resend identical content to an AI for the second time, ask yourself whether it should be cached. Combined with batch processing, which typically offers 50% off non-urgent tasks, you can often cut a bill in half without touching a line of your business logic. That's rare in computing, so it's worth taking advantage of.